Clinical Trial Design • August 25, 2026

Cluster-Randomized Trials: Understanding ICC, Design Effects, and Analysis

Glassmorphic visualization of correlated participants within four randomized trial clusters and a central design-effect prism

A cluster-randomized trial (CRT) randomizes intact groups rather than individuals. Because participants within the same cluster tend to be more alike, the intracluster correlation coefficient (ICC) reduces effective sample size. The resulting design effect inflates the required sample size, while the analysis must account for clustering with cluster-level summaries, mixed-effects models, or generalized estimating equations. The number and size of clusters, unequal cluster sizes, ICC assumptions, and analysis model should be pre-specified.

Cluster randomization is often the only practical way to evaluate an intervention delivered by a hospital, primary-care practice, school, ward, community, or clinician. Randomizing an entire practice may prevent contamination between clinicians; randomizing hospitals may be necessary when an intervention changes staffing or workflow; randomizing communities may be the only feasible route for a population-level program. The design solves an implementation problem, but it creates a statistical one: observations from the same cluster are correlated.

If a trial recruits 1,000 participants from 20 clinics, it does not contain the same amount of information as 1,000 independently randomized participants. Participants in one clinic share staff, protocols, local population characteristics, and care pathways. Their outcomes therefore overlap in predictable ways. Treating them as independent can produce confidence intervals that are too narrow, p-values that are too small, and false-positive conclusions. Correct design and analysis begin with the ICC and end with an estimand that respects the cluster structure.

1. What makes a cluster-randomized trial different?

In an individually randomized trial, the unit of randomization and the unit of analysis are usually the same person. In a CRT, the unit of randomization is the cluster, while the outcome may still be measured for each participant. This distinction affects every stage of the study: recruitment, sample-size planning, allocation, outcome collection, analysis, and interpretation.

The cluster is not merely a nuisance variable. It defines the level at which treatment is assigned and often the level at which the intervention operates. A clinic-level intervention, for example, cannot be randomized independently for two patients treated by the same clinic without a high risk of contamination. The treatment effect is consequently an average contrast between clusters assigned to different arms, often interpreted as the effect of assigning a cluster to an intervention strategy.

CRT estimands should make the population and treatment assignment explicit. The target may be a participant-average effect, a cluster-average effect, or a policy effect that reflects how the intervention performs when rolled out across practices. These effects can differ when cluster sizes vary or when treatment effects are heterogeneous across clusters. The analysis model must therefore be chosen to answer the intended question rather than simply to accommodate correlated data.

2. The intracluster correlation coefficient

The intracluster correlation coefficient, or ICC, quantifies how similar outcomes are for two participants from the same cluster. For a continuous outcome, it is commonly expressed as the between-cluster variance divided by the total variance:

ICC = between-cluster variance / (between-cluster variance + within-cluster variance).

An ICC near zero means that participants in the same cluster are no more alike than participants in different clusters, although clustering can still matter when clusters are large. A higher ICC means that observations within a cluster carry more redundant information. The ICC is a variance component, not a measure of treatment effectiveness, and it should not be interpreted as the proportion of participants who respond similarly.

ICC depends on the outcome, population, cluster definition, and study setting. A hospital-level ICC for mortality may differ substantially from a clinician-level ICC for prescribing behavior or a school-level ICC for a test score. Binary and time-to-event outcomes also require care because the ICC depends on the outcome scale and may vary with the event prevalence. Analysts should use external estimates cautiously and conduct sensitivity analyses over a plausible range.

Reporting the ICC after the trial is useful for future research, but the primary design decision must be made before results are available. The protocol should state the source of the assumed ICC, the outcome scale on which it applies, and the sensitivity range used in the sample-size calculation.

3. Design effect and effective sample size

For equal cluster sizes and a simple individual-level outcome, the conventional design effect is:

Design effect = 1 + (m − 1) × ICC,

where m is the average number of participants per cluster. The individually randomized sample size is multiplied by this factor to obtain the approximate CRT sample size before allowing for attrition. If the average cluster contains 30 participants and the ICC is 0.03, the design effect is 1.87. The trial needs approximately 87% more participants than an individually randomized design with the same nominal power, before other corrections.

The formula makes two planning principles clear. First, even a small ICC can create a substantial inflation when clusters are large. Second, adding participants within existing clusters eventually produces diminishing returns because those participants are correlated with people already recruited. Adding more clusters is often more efficient than making each cluster much larger, especially when the ICC is not negligible.

The effective sample size can be understood as the nominal number of participants divided by the design effect. This is an approximation, not a replacement for a full power calculation, because power also depends on the number of clusters, allocation ratio, cluster-size variation, outcome type, covariate adjustment, and the analysis model. A CRT with very few clusters may have poor small-sample performance even if its nominal participant count looks large.

4. Unequal cluster sizes

Real clusters are rarely equal. A few large hospitals may recruit many more participants than small community practices, and variation in cluster size reduces efficiency. The simple design-effect formula underestimates the required sample size when the cluster-size distribution is highly unequal.

One common approximation uses the coefficient of variation (CV) of cluster size. A modified design effect can be written as approximately:

Design effect for unequal sizes ≈ 1 + [(1 + CV²)m − 1] × ICC.

The exact adjustment depends on the outcome and analysis, but the direction is consistent: greater variation in cluster size generally increases the design effect. Investigators should therefore report how the average cluster size and CV were obtained, not just state a single expected cluster size.

Recruitment targets should account for the possibility that some clusters will contribute fewer participants than planned. Monitoring cluster size during recruitment can help identify whether the trial is losing power. However, changing recruitment or dropping clusters after seeing outcome data can introduce operational and statistical problems; any adaptive rule should be pre-specified and reviewed by the trial statistician.

5. Sample-size planning

A defensible CRT sample-size calculation begins with the individually randomized requirement for the primary outcome and then incorporates clustering. The protocol should state the anticipated effect size, outcome variance or event rate, type I error, target power, allocation ratio, average cluster size, ICC, cluster-size variation, and anticipated loss to follow-up. For binary and survival outcomes, the calculation should use an appropriate CRT-specific method rather than applying the continuous-outcome formula without checking its assumptions.

The number of clusters is often more important than the total number of participants. With few clusters, the treatment effect has limited degrees of freedom, baseline cluster imbalance can be substantial, and standard asymptotic confidence intervals may be unreliable. Small-sample corrections, cluster-level analyses, restricted randomization, or permutation-based inference may be needed. Hayes and Moulton emphasize that the number of clusters, not only the number of individuals, must drive the design decision.

Covariate adjustment can improve precision in a CRT, particularly when cluster-level prognostic variables are known before randomization. The adjusted analysis should include variables used in stratified or restricted randomization and should be specified in the statistical analysis plan. Adjustment can reduce residual between-cluster variation, effectively lowering the design cost, but it does not remove the need to account for within-cluster correlation.

6. Analysis options

Three broad analytical strategies are common. A cluster-level analysis reduces each cluster to a summary, such as a mean, proportion, or log rate, and compares those summaries between arms. It is transparent and naturally respects the randomization unit, but it may lose information and requires careful weighting when cluster sizes differ.

An individual-level mixed-effects model includes a random intercept for cluster and a fixed treatment effect. It can handle continuous, binary, and time-to-event outcomes through suitable generalized or survival extensions. The random-effects structure models the dependence, and the treatment coefficient estimates the effect conditional on the model. Analysts should check whether the model's target effect matches the planned estimand, especially for non-linear outcomes.

Generalized estimating equations (GEE) estimate population-average effects while using a working correlation structure and robust standard errors. In CRTs with a modest number of clusters, the sandwich variance may be biased downward; small-sample corrections such as the Kauermann-Carroll or Fay-Graubard approaches may be required. The choice between mixed models and GEE should be justified by the estimand, cluster count, missing-data assumptions, and outcome distribution.

7. Allocation, contamination, and recruitment

Allocation should be balanced at the cluster level, not only at the participant level. Stratified or restricted randomization can improve balance when the number of clusters is small, but the randomization restrictions must be documented and reflected in the analysis. Cluster concealment is also important: if investigators know the assignment before recruiting participants, recruitment can become differential across arms.

Contamination is a design concern rather than an analytical adjustment. If participants or clinicians in a control cluster receive elements of the intervention, the contrast between arms is diluted. The protocol should define contamination, monitor it when feasible, and describe whether the primary estimand is intention-to-treat assignment or exposure to the intervention. Missing cluster-level outcomes can be especially consequential because losing one cluster may remove many correlated observations and disturb treatment balance.

8. Evidence summary table

Methodology or guidanceContributionPractical implication
Eldridge et al., BMJ 2006Explains the design and analysis implications of cluster randomized trials in primary care.Define the cluster, anticipate ICC, and account for clustering in both design and analysis.
Kerry & Bland, BMJ 2001Introduces the statistical consequences of cluster randomization and sample-size inflation.Use the design effect as a planning concept, not as a justification for ignoring cluster count.
Campbell et al., CONSORT extension 2012Reporting guidance specific to cluster-randomized trials.Report cluster eligibility, recruitment, ICC, cluster flow, and the analysis unit clearly.
Hayes & Moulton, 2017Comprehensive methods for cluster randomized trials, including design, power, and analysis.Plan cluster number, size variation, missingness, and small-sample inference together.
Hemming et al., Trials 2015Guidance for sample-size calculation in cluster randomized trials.State ICC, cluster-size assumptions, coefficient of variation, and inflation for attrition.
CONSORT 2010 extensionMinimum reporting standards for CRT design and results.Make the randomization unit and clustering analysis auditable.
McNeish & Stapleton, 2016Examines multilevel model power when the number of clusters is limited.Use small-sample corrections and avoid relying on nominal participant count alone.
ICH E9(R1)Estimand framework relevant to cluster-level assignment and treatment-policy questions.Define whether the target effect is participant-average, cluster-average, or policy-level.

9. Actionable Steps: Plan and Analyze a Cluster-Randomized Trial

StepResearch actionRequired deliverable
Step 1Define the cluster, the randomization unit, the participant population, the intervention level, and the estimand; document possible contamination pathways.Cluster-level protocol and estimand specification
Step 2Obtain an outcome-specific ICC from reliable external evidence or pilot data and pre-specify a plausible sensitivity range.ICC justification and sensitivity table
Step 3Calculate power using the expected number of clusters, average cluster size, allocation ratio, coefficient of variation, attrition, and the primary analysis model.CRT-specific sample-size report
Step 4Pre-specify restricted or stratified cluster randomization, the primary cluster-aware model, small-sample variance correction, and handling of missing clusters.Randomization and statistical analysis plan
Step 5Report cluster flow, cluster sizes, ICC estimate, treatment effect with uncertainty, contamination, protocol deviations, and sensitivity analyses according to CONSORT.CONSORT-compliant CRT report

10. Common mistakes in CRT research

The most serious error is analyzing participant-level outcomes with an ordinary independent-samples test. That analysis treats correlated observations as independent and usually understates uncertainty. A second error is inflating the sample size with an ICC formula but failing to plan enough clusters; hundreds of participants within a handful of clusters cannot compensate for weak cluster-level replication. A third error is borrowing an ICC from a different outcome or setting without sensitivity analysis. ICC is context-dependent, and an apparently small change can alter the required number of clusters.

Unequal cluster size is another frequent blind spot. Reporting only the mean cluster size hides the loss of efficiency caused by large variation. Finally, some reports fail to state whether the treatment effect is participant-average or cluster-average, or they omit the number of clusters from the primary analysis. A reader cannot evaluate a CRT without knowing how many independent randomized units contributed to the treatment comparison.

11. Relationship to stepped-wedge and other trial designs

A parallel CRT assigns clusters to intervention or control for the study period. A stepped-wedge CRT, covered in an earlier post in this series, assigns clusters to sequences that cross from control to intervention over time. Both designs require attention to ICC, but stepped-wedge trials add within-person or within-cluster serial correlation, period effects, secular trends, and sequence-by-time structure. The simple parallel-trial design effect should not be transferred directly to a stepped-wedge calculation.

Cluster-aware reasoning also connects to multicenter trials, covariate adjustment, and mixed-effects models. The cluster is a source of dependence, a unit of randomization, and often a component of the estimand. Treating those roles separately can produce a model that is mathematically sophisticated but scientifically misaligned. A transparent CRT report makes all three roles explicit.

Researcher's Toolkit: Strengthen Your Cluster Trial Workflow

Cluster-randomized trials require design-effect planning, ICC justification, cluster-aware modeling, and CONSORT-compliant reporting. Lingcore SCI provides specialized tools for medical researchers:

Conclusion

Cluster-randomized trials provide a rigorous solution when interventions operate at the level of practices, hospitals, schools, or communities, but the design changes what counts as information. The ICC measures within-cluster similarity; the design effect translates that similarity into sample-size inflation; unequal cluster sizes and a limited number of clusters determine whether the nominal calculation is credible; and cluster-aware models protect inference after recruitment. A strong CRT protocol defines the randomization unit and estimand, justifies the ICC, plans the number and size of clusters, pre-specifies the analysis, and reports cluster flow and uncertainty transparently. These steps turn a practical implementation design into defensible clinical evidence.