Network Meta-Analysis: A Practical Guide

Written and reviewed by Andrew Collins, PhD May 12, 2026 9 min read

A network meta-analysis is a method for comparing three or more treatments at once by combining direct evidence, where two treatments were tested head to head in the same trial, with indirect evidence, where two treatments were each compared against a shared common comparator but never against each other. Instead of running a separate pairwise analysis for every pair, it knits all the eligible studies into a single connected evidence network in which treatments are nodes and the trials that compared them are the edges. The payoff is a coherent ranking of every option on the same scale, even for pairs that no trial has ever compared directly, which is why it has become the standard tool for clinical guidelines and reimbursement decisions where many alternatives compete.

How indirect evidence is borrowed across the network

The engine behind a network meta-analysis is the indirect comparison. Suppose no trial ever compared treatment A against treatment B, but several trials compared A against a common comparator C, and others compared B against C. If A reduces an outcome relative to C by a certain amount, and B reduces the same outcome relative to C by a smaller amount, then the difference between those two contrasts estimates how A would perform against B. In effect, the shared comparator C acts as a bridge that lets the analysis carry information from one corner of the network to another. This borrowing of strength is what allows the method to fill in pairwise comparisons that the trial evidence never measured.

When a pair of treatments has both a head to head trial and an indirect path through a common comparator, the network meta-analysis combines the two into a single mixed estimate, weighting each source by its precision. Because every contrast is expressed on the same scale, usually as a log odds ratio, log risk ratio, or a mean difference, the whole network is fitted as one statistical model rather than as a pile of disconnected comparisons. The result is a full matrix of relative effects between every pair of treatments, each with its own confidence interval or credible interval.

The transitivity assumption

Indirect comparison is only valid if the studies are similar enough to be linked. This is the transitivity assumption: the trials comparing A with C must be clinically and methodologically comparable to the trials comparing B with C, so that the common comparator C behaves the same way across both sets. If the A versus C trials enrolled younger, healthier patients while the B versus C trials enrolled older, sicker ones, then the comparator is not really the same anchor in both places, and the borrowed estimate of A versus B will be biased. Transitivity is a judgement about the effect modifiers, the patient and design characteristics that change how well a treatment works, and it cannot be proven by the data alone. You defend it by tabulating dose, follow-up, baseline risk, and population across the network and arguing that they are balanced.

The consistency assumption

Transitivity has a measurable counterpart called consistency, sometimes named coherence. It states that the direct evidence and the indirect evidence for the same pair should agree. Where a closed loop exists in the network, for example A versus B, B versus C, and A versus C all measured, you can compare the direct estimate of a contrast against the indirect estimate implied by the rest of the loop. A large, unexplained gap signals inconsistency, which usually means transitivity was violated somewhere. Analysts test this with node-splitting, which separates direct from indirect contributions at each node, or with a design by treatment interaction model that checks the whole network at once. Consistency is the statistical alarm; transitivity is the conceptual cause.

Ranking treatments and the SUCRA caveat

Because a network meta-analysis estimates every pairwise contrast, it can also order the treatments from best to worst. The most common summary is the surface under the cumulative ranking curve, abbreviated SUCRA, a single number between 0 and 100 percent that captures the probability a treatment ranks among the better options. A related quantity is the probability of being best, the chance that a given treatment is the top performer. These rankings are seductive and widely misread. A treatment can hold the highest SUCRA value while its advantage over the runner up is tiny and statistically uncertain, so a ranking table must always be read next to the actual effect sizes and their intervals, never on its own. Rankings are also unstable when a node rests on a single small trial, and the probability of being best is sensitive to how many treatments crowd the top of the network. Treat the ranking as a hypothesis-ordering device, not a verdict.

Frequentist and Bayesian approaches

There are two main ways to fit the model. The frequentist approach, often implemented with multivariate or graph-theoretical methods, produces relative effects, confidence intervals, and SUCRA values without specifying prior beliefs, and it tends to run quickly. The Bayesian approach fits the same network through Markov chain Monte Carlo sampling, placing prior distributions on the treatment effects and on the between-study heterogeneity. Its strengths are that it handles sparse networks and multi-arm trials naturally, propagates all uncertainty into the rankings, and yields intuitive probability statements such as the chance a treatment is best. In practice the two methods give very similar point estimates when the data are informative; the choice is usually driven by the complexity of the network and the analyst's software fluency. Both rest on the same transitivity and consistency assumptions, so neither rescues a poorly connected network.

The main pitfalls

Several traps recur often enough that reviewers look for them by name. The first is a sparse or poorly connected network, where some treatments hang on a single thin edge, leaving estimates that swing wildly. The second is unexamined heterogeneity and inconsistency, reported without the loop-specific checks that would reveal them. The third is over-reading the rankings, presenting a SUCRA league table as if it settled the clinical question. The fourth is a failure of transitivity hidden by silence, where the authors never show that the trials feeding each comparison were comparable.

  • Map the network first and disclose it, so readers can see which contrasts rest on direct trials and which are purely indirect.
  • Defend transitivity explicitly by tabulating the effect modifiers across the comparisons rather than asserting comparability.
  • Test consistency in every closed loop and report what the checks found, even when they are reassuring.
  • Pair rankings with effect estimates and their intervals, and never let a SUCRA value stand alone as the conclusion.

The groundwork is the same systematic search and data extraction that any review demands, so the distinction between a protocol-driven review and the quantitative synthesis that follows applies here too. If you are new to pooling effects across studies, our primer on how a meta-analysis combines results into one estimate is the right starting point, and the decision between a single shared effect or a distribution of effects across studies carries directly into how the network model handles heterogeneity. For an ordinary pairwise comparison you can pool event counts in our browser-based effect-size calculator, and once you have a single contrast you can visualise the pooled result as a publication-ready forest plot before scaling up to a full network.

Frequently asked questions

How do you write a network meta-analysis?
Start with a registered protocol and a systematic search, then extract every eligible comparison and map the evidence network. Choose an effect measure, fit a frequentist or Bayesian model, and report relative effects for all treatment pairs. Crucially, justify the transitivity assumption, test consistency in each loop, and present rankings alongside effect estimates. Follow the PRISMA extension for network meta-analysis when reporting.
How many studies do you need for a network meta-analysis?
There is no fixed minimum; what matters is a connected network where every treatment links to the others through at least one path. A technically valid analysis can run on a handful of trials, but estimates from a single thin edge are fragile. Aim for several trials per comparison so that heterogeneity and consistency can be assessed and the rankings remain stable rather than driven by one small study.
What are the disadvantages of network meta-analysis?
It rests on the transitivity assumption, which cannot be proven and is easily violated when populations or designs differ across comparisons. Sparse networks produce unstable estimates, inconsistency between direct and indirect evidence can bias results, and treatment rankings are routinely over-interpreted. It is also more complex to fit and report than a pairwise analysis, demanding more methodological expertise to defend at peer review.
What is the difference between direct and indirect comparison?
A direct comparison comes from trials that tested two treatments head to head in the same study. An indirect comparison estimates the contrast between two treatments that were never compared directly, by linking them through a shared common comparator: if A beats C and B beats C, the difference estimates A versus B. Network meta-analysis combines both into a single mixed estimate when both exist.

Written and reviewed by

Andrew Collins, PhD

Independent Research Advisor

Andrew Collins is an independent research advisor with extensive experience in systematic review methodology and quantitative evidence synthesis. He has assisted researchers with protocol development, literature screening, data extraction, and meta-analysis, helping them navigate each stage of the review process. Andrew is committed to promoting rigorous and transparent research practices that support evidence-based decision making.

The methods in this guide follow the Cochrane Handbook and Borenstein and colleagues' Introduction to Meta-Analysis, and the statistics behind our tools are validated against the metafor package in R and statsmodels in Python.