A network meta-analysis is a method for comparing three or more treatments at once by combining direct evidence, where two treatments were tested head to head in the same trial, with indirect evidence, where two treatments were each compared against a shared common comparator but never against each other. Instead of running a separate pairwise analysis for every pair, it knits all the eligible studies into a single connected evidence network in which treatments are nodes and the trials that compared them are the edges. The payoff is a coherent ranking of every option on the same scale, even for pairs that no trial has ever compared directly, which is why it has become the standard tool for clinical guidelines and reimbursement decisions where many alternatives compete.
How indirect evidence is borrowed across the network
The engine behind a network meta-analysis is the indirect comparison. Suppose no trial ever compared treatment A against treatment B, but several trials compared A against a common comparator C, and others compared B against C. If A reduces an outcome relative to C by a certain amount, and B reduces the same outcome relative to C by a smaller amount, then the difference between those two contrasts estimates how A would perform against B. In effect, the shared comparator C acts as a bridge that lets the analysis carry information from one corner of the network to another. This borrowing of strength is what allows the method to fill in pairwise comparisons that the trial evidence never measured.
When a pair of treatments has both a head to head trial and an indirect path through a common comparator, the network meta-analysis combines the two into a single mixed estimate, weighting each source by its precision. Because every contrast is expressed on the same scale, usually as a log odds ratio, log risk ratio, or a mean difference, the whole network is fitted as one statistical model rather than as a pile of disconnected comparisons. The result is a full matrix of relative effects between every pair of treatments, each with its own confidence interval or credible interval.
The transitivity assumption
Indirect comparison is only valid if the studies are similar enough to be linked. This is the transitivity assumption: the trials comparing A with C must be clinically and methodologically comparable to the trials comparing B with C, so that the common comparator C behaves the same way across both sets. If the A versus C trials enrolled younger, healthier patients while the B versus C trials enrolled older, sicker ones, then the comparator is not really the same anchor in both places, and the borrowed estimate of A versus B will be biased. Transitivity is a judgement about the effect modifiers, the patient and design characteristics that change how well a treatment works, and it cannot be proven by the data alone. You defend it by tabulating dose, follow-up, baseline risk, and population across the network and arguing that they are balanced.
The consistency assumption
Transitivity has a measurable counterpart called consistency, sometimes named coherence. It states that the direct evidence and the indirect evidence for the same pair should agree. Where a closed loop exists in the network, for example A versus B, B versus C, and A versus C all measured, you can compare the direct estimate of a contrast against the indirect estimate implied by the rest of the loop. A large, unexplained gap signals inconsistency, which usually means transitivity was violated somewhere. Analysts test this with node-splitting, which separates direct from indirect contributions at each node, or with a design by treatment interaction model that checks the whole network at once. Consistency is the statistical alarm; transitivity is the conceptual cause.
Ranking treatments and the SUCRA caveat
Because a network meta-analysis estimates every pairwise contrast, it can also order the treatments from best to worst. The most common summary is the surface under the cumulative ranking curve, abbreviated SUCRA, a single number between 0 and 100 percent that captures the probability a treatment ranks among the better options. A related quantity is the probability of being best, the chance that a given treatment is the top performer. These rankings are seductive and widely misread. A treatment can hold the highest SUCRA value while its advantage over the runner up is tiny and statistically uncertain, so a ranking table must always be read next to the actual effect sizes and their intervals, never on its own. Rankings are also unstable when a node rests on a single small trial, and the probability of being best is sensitive to how many treatments crowd the top of the network. Treat the ranking as a hypothesis-ordering device, not a verdict.
Frequentist and Bayesian approaches
There are two main ways to fit the model. The frequentist approach, often implemented with multivariate or graph-theoretical methods, produces relative effects, confidence intervals, and SUCRA values without specifying prior beliefs, and it tends to run quickly. The Bayesian approach fits the same network through Markov chain Monte Carlo sampling, placing prior distributions on the treatment effects and on the between-study heterogeneity. Its strengths are that it handles sparse networks and multi-arm trials naturally, propagates all uncertainty into the rankings, and yields intuitive probability statements such as the chance a treatment is best. In practice the two methods give very similar point estimates when the data are informative; the choice is usually driven by the complexity of the network and the analyst's software fluency. Both rest on the same transitivity and consistency assumptions, so neither rescues a poorly connected network.
The main pitfalls
Several traps recur often enough that reviewers look for them by name. The first is a sparse or poorly connected network, where some treatments hang on a single thin edge, leaving estimates that swing wildly. The second is unexamined heterogeneity and inconsistency, reported without the loop-specific checks that would reveal them. The third is over-reading the rankings, presenting a SUCRA league table as if it settled the clinical question. The fourth is a failure of transitivity hidden by silence, where the authors never show that the trials feeding each comparison were comparable.
- Map the network first and disclose it, so readers can see which contrasts rest on direct trials and which are purely indirect.
- Defend transitivity explicitly by tabulating the effect modifiers across the comparisons rather than asserting comparability.
- Test consistency in every closed loop and report what the checks found, even when they are reassuring.
- Pair rankings with effect estimates and their intervals, and never let a SUCRA value stand alone as the conclusion.
The groundwork is the same systematic search and data extraction that any review demands, so the distinction between a protocol-driven review and the quantitative synthesis that follows applies here too. If you are new to pooling effects across studies, our primer on how a meta-analysis combines results into one estimate is the right starting point, and the decision between a single shared effect or a distribution of effects across studies carries directly into how the network model handles heterogeneity. For an ordinary pairwise comparison you can pool event counts in our browser-based effect-size calculator, and once you have a single contrast you can visualise the pooled result as a publication-ready forest plot before scaling up to a full network.