Statistical heterogeneity is the variation in true effects across the studies in a meta-analysis that goes beyond what sampling error alone would produce, and I-squared is the most widely reported way to quantify it. I-squared interpretation tells you what percentage of the total observed variation in effect estimates is due to genuine differences between studies rather than chance. An I-squared of zero means the studies look like repeated samples of one underlying effect, while a large I-squared means the effect itself shifts from study to study. The number does not tell you whether that variation is good or bad, only how much of it exists, so it must always be read alongside the size, direction, and clinical meaning of the pooled result on your forest plot reading guide.
Where heterogeneity comes from
Studies that address the same question are rarely identical. They recruit different populations, apply slightly different doses or comparators, measure outcomes at different time points, and carry different risks of bias. Each of these design differences can move the true underlying effect, producing what methodologists call clinical and methodological diversity. When that diversity translates into differences in the measured effect sizes, it shows up as between-study variance. This is the substantive signal that I-squared and its companion statistics are trying to capture. Before reaching for any number, it is worth listing the plausible sources of variation in your own set of studies, because the statistics describe heterogeneity but never explain it. Understanding how pooling combines study evidence makes clear why this variation matters: pooling assumes the studies are estimating something comparable.
Cochran's Q and why it is only a starting point
The classical test for heterogeneity is Cochran's Q. It is a weighted sum of the squared deviations of each study's effect estimate from the pooled fixed effect estimate, with each study weighted by the inverse of its variance. Under the null hypothesis that every study shares one common true effect, Q follows a chi-squared distribution with degrees of freedom equal to the number of studies minus one. A significant Q is evidence that the effects differ by more than chance.
Cochran's Q has two well known weaknesses that shape how you should use it. With only a handful of studies it has low statistical power, so a non-significant Q does not prove the effects are homogeneous, it may simply mean there were too few studies to detect real differences. Conversely, when a meta-analysis includes many large studies, Q becomes oversensitive and flags trivial, clinically meaningless differences as statistically significant. Because Q depends so heavily on the number and size of studies, it works poorly as a standalone descriptor, which is exactly the gap that I-squared was designed to fill.
How I-squared is derived
I-squared rescales Cochran's Q into a quantity that does not depend on the number of studies. The formula is I-squared equals Q minus the degrees of freedom, all divided by Q, expressed as a percentage, with any negative result set to zero. In words, I-squared is the proportion of the total variability in effect estimates that is attributable to true between-study heterogeneity rather than to within-study sampling error. Because it is a ratio rather than an absolute amount, I-squared lets you compare the consistency of meta-analyses that use different outcomes and scales, which is why it is the figure most journals expect to see reported.
A point that trips up many researchers is that I-squared is relative, not absolute. It describes the share of variation that is real, so a meta-analysis of very precise studies can post a high I-squared even when the actual spread of effects is small, simply because so little of the variation is noise. This is one reason a high I-squared on its own does not condemn a synthesis.
The conventional thresholds and the Cochrane caveat
The Cochrane Handbook offers rough benchmarks for I-squared interpretation, while stressing that they are a guide and not a rule:
- Around 25 percent is often described as low heterogeneity, where the studies are reasonably consistent.
- Around 50 percent is treated as moderate heterogeneity, a signal to investigate rather than ignore.
- Around 75 percent is regarded as high or considerable heterogeneity, where a single pooled estimate may obscure important differences.
Cochrane is explicit that these cut points are context-dependent. The same I-squared value can be acceptable in one field and alarming in another, because what matters is whether the variation changes the practical conclusion. The importance of an I-squared value depends on the magnitude and direction of effects and on the strength of the evidence for heterogeneity, such as the Q test p-value and the confidence interval around I-squared itself, which is often wide when there are few studies. Treating 50 percent as a hard pass or fail line is a misuse of the statistic.
Tau-squared and H-squared: the other two numbers
I-squared is rarely the whole story. Tau-squared is the estimated absolute between-study variance, the variance of the distribution of true effects on the scale of the outcome. Its square root, tau, is the between-study standard deviation and is the most interpretable measure of how far apart the true effects are. Where I-squared tells you the proportion of variation that is real, tau-squared tells you how big that real variation actually is, which is why it drives the width of a prediction interval for future studies.
H-squared is a third related index, defined as Q divided by its degrees of freedom. It equals one when there is no heterogeneity and grows above one as heterogeneity increases, so it describes the ratio of the total variation to the variation expected from sampling error alone. Reporting tau-squared and H-squared alongside I-squared gives readers a fuller picture than any single number, because together they separate the proportion, the absolute size, and the scale of the heterogeneity.
What to do about high heterogeneity
A large I-squared is an invitation to investigate, not a verdict. The standard toolkit moves from modelling choices to active exploration of the causes.
Choose the right model
When meaningful heterogeneity is present, a random-effects model is usually the honest default, because it assumes the true effects are drawn from a distribution and widens the confidence interval to reflect that uncertainty. A fixed-effect model assumes one shared true effect and will understate uncertainty when heterogeneity is real. The trade-off between these two assumptions is laid out in our guide to fixed-effect versus random-effects pooling, and it is the first decision to revisit when I-squared climbs.
Explore the sources
- Subgroup analysis splits the studies by a prespecified characteristic, such as dose, population, or risk of bias, to see whether the effect differs across groups.
- Meta-regression models the effect size as a function of one or more study-level covariates, quantifying how much of the heterogeneity a moderator explains.
- Sensitivity analysis re-runs the synthesis under different reasonable choices, for example excluding high risk of bias studies, to test how robust the pooled estimate is.
- Outlier and influence checks identify single studies that distort the pooled effect, which you can examine with leave-one-out diagnostics in our meta-analysis calculator.
Prespecify these analyses in your protocol wherever possible. Subgroup comparisons and meta-regressions decided after seeing the data are prone to false positives, and reviewers know to discount them.
Why high I-squared does not invalidate a meta-analysis
A common misconception is that a high I-squared means the studies should never have been combined. That is not so. Heterogeneity is a property to be characterised and explained, not a flaw that disqualifies synthesis. A random-effects estimate with a wide but honest confidence interval, paired with a clear account of why the effects vary, can be far more informative than a falsely precise fixed-effect number. The real failure is not high I-squared, it is failing to report it, failing to investigate it, or pooling studies that are genuinely answering different questions. Used well, the combination of I-squared, tau-squared, and a transparent exploration of moderators turns heterogeneity from a threat into one of the most useful outputs of your review. When you are ready to visualise the spread of your effects, plot your effect sizes and read the per-study estimates against the pooled line.