What Is a Meta-Analysis? A Researcher’s Guide

Written and reviewed by Rachel Bennett, MPH May 29, 2026 9 min read

A meta-analysis is a statistical method that combines the numerical results of several independent studies that ask the same research question, producing a single pooled effect estimate with a confidence interval that is more precise than any individual study could offer on its own. Rather than counting how many studies were positive or negative, a meta-analysis weights every study by how much information it carries and synthesises those weighted results into one quantitative answer. It is the analytical engine that usually sits at the heart of a systematic review, and its output is most often displayed in a visual summary called a forest plot.

How a meta-analysis differs from a systematic review

The two terms are frequently treated as interchangeable, but they describe different things. A systematic review is the full process of locating, appraising, and summarising all the evidence that meets a predefined protocol, and it can be entirely qualitative. A meta-analysis is the optional quantitative synthesis step that some systematic reviews include when the included studies are similar enough to pool. In other words, every meta-analysis worth trusting grows out of a systematic search, but not every systematic review ends in a meta-analysis. A narrative review, by contrast, summarises evidence in prose without a reproducible protocol or pooled statistics, which makes it more prone to selective reporting and author bias.

The practical consequence is that the credibility of a pooled estimate depends entirely on the rigour of the review that feeds it. If the search was incomplete or the screening was sloppy, the most elegant pooling model in the world will still return a biased number.

The steps of a meta-analysis from question to interpretation

A defensible quantitative synthesis follows a disciplined sequence. Each step is documented so that a second researcher could reproduce the result from the same data.

  • Define a focused question using a structured framework such as population, intervention, comparator, and outcome, so the eligibility rules are unambiguous before any searching begins.
  • Search systematically across multiple databases with a documented strategy, supplemented by reference checking and grey literature, to reduce the risk of missing relevant evidence.
  • Screen and select studies against the protocol, ideally with two independent reviewers, and record the flow of records in a PRISMA diagram.
  • Extract data and effect sizes, capturing the event counts, means, standard deviations, or sample sizes needed to compute a comparable number from each study.
  • Assess risk of bias in each included study, because a precise pooled estimate built on flawed trials is precise but wrong.
  • Pool the effect sizes with an appropriate statistical model, then examine heterogeneity, publication bias, and the robustness of the finding.

Choosing an effect measure

Before anything can be pooled, every study must be expressed on a common scale called the effect measure. For binary outcomes such as recovered or not recovered, researchers commonly use the odds ratio, the risk ratio, or the risk difference. For continuous outcomes such as a symptom score, the mean difference is used when studies share the same scale, and the standardised mean difference is used when they do not. Ratio measures are pooled on the logarithmic scale so that the sampling distribution is roughly symmetric, then converted back for reporting. If you are unsure which relative measure to report, the comparison of an odds ratio against a risk ratio is a useful starting point, and you can compute a single combined value quickly with the summary effect calculator.

Weighting studies by precision

The defining feature of a meta-analysis is that studies are not treated equally. Each study is given a study weight that reflects how much information it carries, which is driven by its sample size and the variance of its estimate. In the standard inverse-variance method, a study with a tight, precise estimate receives a large weight, while a small noisy study receives a small one. This is why a single large, well-conducted trial can dominate a pooled result, and why adding several tiny studies often barely moves the combined number.

Fixed-effect versus random-effects models

Two statistical models dominate the pooling step, and the choice changes both the result and its interpretation. A fixed-effect model assumes every study estimates one identical true effect and that the only reason estimates differ is sampling error, so it pools toward that single value. A random-effects model instead assumes the true effect varies from study to study around an average, so it estimates that average while widening the confidence interval to acknowledge the between-study variation. Because real-world interventions and populations rarely produce one identical effect, the random-effects model is the more common default, though it gives smaller studies relatively more influence. The full trade-off is worked through in our guide on choosing between fixed and random effects.

Assessing heterogeneity

Heterogeneity describes how much the true effects differ across studies beyond what chance alone would explain. The most cited summary is the I-squared statistic, which expresses the percentage of total variation attributable to genuine between-study differences rather than sampling error, supported by the Cochran Q test and the tau-squared estimate of between-study variance. High I-squared values warn that a single pooled number may mask important differences, prompting subgroup analysis, meta-regression, or a decision not to pool at all. For a deeper treatment, see how to interpret the heterogeneity statistics.

Checking for publication bias

Studies with statistically significant or favourable results are more likely to be published, which can inflate a pooled estimate. To probe this, researchers inspect a funnel plot, where each study is plotted against its precision; a symmetric inverted funnel suggests little bias, while a gap in one corner hints that smaller null studies are missing. Asymmetry can be tested formally and explored with trim-and-fill or related sensitivity methods. You can generate the diagnostic yourself with the funnel plot tool, and the broader comparison of what a forest plot shows versus a funnel plot clarifies why both belong in a complete report.

Visualising the result in a forest plot

The headline output of a meta-analysis is the forest plot. Each included study appears as a marker whose size reflects its weight, with a horizontal line for its confidence interval, and the pooled estimate sits at the bottom as a diamond whose width is the combined confidence interval. A reference line of no effect lets a reader judge at a glance which studies favoured the intervention and whether the synthesis as a whole crossed into significance. You can assemble one from your own extracted numbers in minutes with the browser-based plotting tool.

When a meta-analysis is and is not appropriate

Pooling is powerful only when the underlying studies are genuinely comparable. A meta-analysis is appropriate when the included studies share a coherent question, a consistent effect measure, and a defensible level of clinical and methodological similarity. It is not appropriate, and can be actively misleading, when the interventions or populations are too heterogeneous to represent the same underlying effect, when outcomes are measured in incompatible ways, or when there are simply too few studies for the heterogeneity estimate to be stable. Pooling two or three wildly different trials produces a number that looks authoritative while describing nothing real. In those cases a structured narrative synthesis, or pooling only within sensible subgroups, is the honest choice.

Used well, a meta-analysis turns a scattered literature into a single, weighted, transparent estimate that readers can scrutinise and reproduce. Used carelessly, it lends false certainty to evidence that was never coherent enough to combine. The craft lies in knowing the difference before you press the pool button.

Frequently asked questions

What is the difference between a meta-analysis and a systematic review?
A systematic review is the full reproducible process of finding, appraising, and summarising all evidence that fits a protocol, and it can be purely qualitative. A meta-analysis is the optional statistical step that pools the numerical results when the included studies are similar enough. Every sound meta-analysis sits inside a systematic search, but many systematic reviews do not include a meta-analysis.
How many studies do you need for a meta-analysis?
There is no strict minimum, and two studies can technically be combined, but estimates of heterogeneity become unstable with very few studies. Many methodologists are cautious about pooling fewer than five studies, especially with a random-effects model, because the between-study variance cannot be estimated reliably. When studies are scarce or too different, a structured narrative synthesis is often the more honest choice.
Should I use a fixed-effect or random-effects model?
Use a fixed-effect model only when it is reasonable to assume every study estimates one identical true effect. Because real interventions and populations usually vary, the random-effects model is the more common default, since it allows the true effect to differ across studies and widens the confidence interval accordingly. Always pre-specify the model in your protocol rather than choosing it after seeing the results.

Written and reviewed by

Rachel Bennett, MPH

Evidence Synthesis Consultant

Rachel Bennett is an evidence synthesis consultant who helps researchers conduct systematic reviews and meta-analyses in healthcare and social sciences. Her work includes designing search strategies, evaluating study quality, synthesizing findings, and preparing manuscripts for publication. Rachel is passionate about making research accessible and ensuring that evidence is presented clearly and accurately.

The methods in this guide follow the Cochrane Handbook and Borenstein and colleagues' Introduction to Meta-Analysis, and the statistics behind our tools are validated against the metafor package in R and statsmodels in Python.