Undue reliance on I2 in assessing heterogeneity may mislead

Gerta Rücker, Guido Schwarzer, James R. Carpenter, Martin SchumacherView original
OverviewBalancedalloy voice
Imagine a meta-analysis lands on a journal editor's desk. Seventy trials, thousands of patients, and a clear pooled result. The researchers check their heterogeneity statistic, I-squared, which is the standard measure, and it reads eighteen percent. That's low. The studies are consistent. They pool the results, the paper gets published, and the finding shapes a clinical guideline. Now imagine that same body of evidence, with the same true between-study differences and the same underlying biology, but the trials were each run with four times as many patients. The I-squared doesn't read eighteen percent anymore. It reads nearly thirty percent. If the trials had sixteen times the patients, I-squared climbs to eighty-five percent. With sixty-four times the patients, it approaches ninety-six percent — a number the Cochrane Handbook would call considerable heterogeneity. Nothing about the actual differences between studies changed. Only the sample sizes did. That is the problem Rücker, Schwarzer, Carpenter, and Schumacher set out to document, and it has been sitting quietly inside thousands of published meta-analyses. To understand why it matters, start with what a meta-analysis is trying to do. When you combine results from multiple studies, you get a more precise estimate of a treatment effect than any single trial could deliver. But studies differ — in their patients, their designs, and their measurements — and some of that difference shows up in the numbers themselves. Rücker and colleagues focus on statistical heterogeneity, which is the variability in treatment effect estimates that you can actually quantify on the outcome scale. This is the variation that makes researchers question whether it is defensible to average across studies at all. If the true treatment effect differs dramatically from trial to trial, a pooled average may not represent any real clinical situation. So the heterogeneity assessment is not just a formality. It is the gate that controls whether the pooled number gets used. The tool most reviewers use to pass through that gate is I-squared. It derives from Cochran's Q statistic, which is a weighted sum of squared deviations — each study's estimated effect compared to the pooled estimate, squared, and multiplied by that study's weight. I-squared takes Q, subtracts the degrees of freedom, which is the number of studies minus one, divides that difference by Q, and expresses the result as a percentage. The interpretation sounds intuitive: it is the proportion of total variability in the estimates that comes from real between-study differences rather than sampling error. The Cochrane Handbook offers rough thresholds: zero to forty percent might not be important, thirty to sixty may be moderate, fifty to ninety may be substantial, and seventy-five to one hundred is considerable. Many reviewers in practice use fifty percent as a binary cutoff. Here is the mechanical problem. Under a random-effects model, the total variance for any given study is the sum of two things: the within-study sampling variance, which shrinks as sample size grows, and the between-study variance, tau-squared, which is a fixed property of how much the true treatment effects differ across studies. I-squared is a ratio of the between-study component to the total. So as within-study variance falls, meaning studies get larger and more precise, the between-study component becomes a larger share of the total, even if tau-squared itself hasn't moved. The percentage climbs not because heterogeneity grew, but because sampling error shrank. Rücker and colleagues demonstrate this with a clean simulation. They took a real seventy-trial meta-analysis of thrombolytic therapy, which are clot-dissolving drugs used in acute myocardial infarction. The original data gave a DerSimonian-Laird estimate of tau-squared of 0.018, an I-squared of eighteen point six percent with a confidence interval spanning zero to forty, and a Q statistic of eighty-five with a p-value of 0.095. There was no statistical evidence of heterogeneity. Then they inflated within-trial precision by integer multiplication factors, drawing new effect estimates from the same random-effects model but with sampling variance divided by M. The between-study variance stayed fixed. Only the precision of each trial increased. The results are stark. At an inflation factor of four, tau-squared was estimated at 0.008 — essentially the same, just random fluctuation — while I-squared rose to twenty-nine point two percent and Q climbed to ninety-eight with a p-value now crossing the significance threshold at 0.014. At a factor of sixteen, tau-squared was 0.027, still in the same range, but I-squared had reached eighty-four point eight percent and Q was four hundred fifty-four. At a factor of sixty-four, tau-squared was 0.028 — indistinguishable from the original — and I-squared was ninety-six percent. Q had reached one thousand seven hundred eight. The statistic that was supposed to measure between-study differences was instead tracking how large the trials were. The authors then checked whether this pattern holds across real meta-analyses, not just in a single simulated one. In a sample of one hundred fifty-seven binary-endpoint meta-analyses, they fitted a linear model predicting I-squared from two variables: tau-squared and the logarithm of the median study size. Both were strongly significant. The coefficient for logarithm of median study size was eight point five — meaning each doubling of typical study size pushed I-squared up by about six points — and the model explained sixty-six percent of the variance in I-squared across those one hundred fifty-seven analyses. Tau-squared matters too, but the point is that study size matters independently, and the field has generally not accounted for that. There is a further wrinkle Rücker and colleagues note. If you have a large trial whose estimated effect happens to be close to the pooled average, removing it is one of the most efficient ways to reduce I-squared. That should give pause. A statistic that goes down when you exclude a well-powered, on-average trial is not straightforwardly measuring the heterogeneity of the evidence base. So what should researchers use instead? Tau-squared. The between-study variance lives on the outcome scale — the same scale as the treatment effects themselves. That means clinicians can ask a clinically meaningful question: is the spread of true effects across studies wide enough to change a treatment decision? Rücker and colleagues give a concrete anchor for odds-ratio outcomes. If the true effects across studies ranged from an odds ratio of 0.8 to 1.0 to 1.25, those values on the log scale correspond to tau — the standard deviation of true effects — of 0.22, which means tau-squared of roughly 0.05. A reviewer can look at their estimated tau-squared, compare it to that threshold, and ask whether the spread matters clinically. Crucially, that comparison doesn't change just because the trials enrolled more patients. The DerSimonian-Laird estimator is the standard tool for estimating tau-squared, and it is what Rücker and colleagues use throughout. They are candid that tau-squared has its own uncertainty — in the thrombolysis simulations, it fluctuated noticeably across runs, and the literature on heterogeneity estimation acknowledges that with few studies, any variance estimate will be imprecise. But imprecision is a reason to interpret carefully, not a reason to prefer a measure that is systematically biased by study size. An uncertain answer to the right question beats a confident answer to the wrong one. The implications reach into the infrastructure of evidence-based medicine. I-squared appears in almost every published systematic review. It is used by guideline committees, by journal reviewers deciding whether a meta-analysis is valid, and by clinicians trying to understand whether a pooled result applies to their patient. If large, well-powered trials routinely push I-squared toward alarming values even when the actual between-study differences are modest, the field may be declining to pool evidence that should be pooled — or raising false alarms that send researchers hunting for moderators that don't exist. The practical message from Rücker and colleagues is direct: when you read a meta-analysis, find the tau-squared, not the I-squared. Ask what the estimated spread of true effects is and whether that spread is large enough to change what you would do clinically. If tau-squared is small relative to a meaningful clinical difference, pooling is defensible — regardless of whether I-squared reads thirty percent or ninety percent. The percentage that has come to dominate heterogeneity assessment is telling you something, but it is not telling you what most readers think it is. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Imagine a meta-analysis lands on a journal editor's desk. Seventy trials, thousands of patients, and a clear pooled result. The researchers check their heterogeneity statistic, I-squared, which is the standard measure, and it reads eighteen percent.

That's low. The studies are consistent. They pool the results, the paper gets published, and the finding shapes a clinical guideline.

Now imagine that same body of evidence, with the same true between-study differences and the same underlying biology, but the trials were each run with four times as many patients. The I-squared doesn't read eighteen percent anymore. It reads nearly thirty percent.

If the trials had sixteen times the patients, I-squared climbs to eighty-five percent. With sixty-four times the patients, it approaches ninety-six percent — a number the Cochrane Handbook would call considerable heterogeneity. Nothing about the actual differences between studies changed.

Only the sample sizes did. That is the problem Rücker, Schwarzer, Carpenter, and Schumacher set out to document, and it has been sitting quietly inside thousands of published meta-analyses.

To understand why it matters, start with what a meta-analysis is trying to do. When you combine results from multiple studies, you get a more precise estimate of a treatment effect than any single trial could deliver. But studies differ — in their patients, their designs, and their measurements — and some of that difference shows up in the numbers themselves.

Rücker and colleagues focus on statistical heterogeneity, which is the variability in treatment effect estimates that you can actually quantify on the outcome scale. This is the variation that makes researchers question whether it is defensible to average across studies at all. If the true treatment effect differs dramatically from trial to trial, a pooled average may not represent any real clinical situation.

So the heterogeneity assessment is not just a formality. It is the gate that controls whether the pooled number gets used.

The tool most reviewers use to pass through that gate is I-squared. It derives from Cochran's Q statistic, which is a weighted sum of squared deviations — each study's estimated effect compared to the pooled estimate, squared, and multiplied by that study's weight. I-squared takes Q, subtracts the degrees of freedom, which is the number of studies minus one, divides that difference by Q, and expresses the result as a percentage.

The interpretation sounds intuitive: it is the proportion of total variability in the estimates that comes from real between-study differences rather than sampling error. The Cochrane Handbook offers rough thresholds: zero to forty percent might not be important, thirty to sixty may be moderate, fifty to ninety may be substantial, and seventy-five to one hundred is considerable. Many reviewers in practice use fifty percent as a binary cutoff.

Here is the mechanical problem. Under a random-effects model, the total variance for any given study is the sum of two things: the within-study sampling variance, which shrinks as sample size grows, and the between-study variance, tau-squared, which is a fixed property of how much the true treatment effects differ across studies. I-squared is a ratio of the between-study component to the total.

So as within-study variance falls, meaning studies get larger and more precise, the between-study component becomes a larger share of the total, even if tau-squared itself hasn't moved. The percentage climbs not because heterogeneity grew, but because sampling error shrank.

Rücker and colleagues demonstrate this with a clean simulation. They took a real seventy-trial meta-analysis of thrombolytic therapy, which are clot-dissolving drugs used in acute myocardial infarction. The original data gave a DerSimonian-Laird estimate of tau-squared of 0.018, an I-squared of eighteen point six percent with a confidence interval spanning zero to forty, and a Q statistic of eighty-five with a p-value of 0.095.

There was no statistical evidence of heterogeneity. Then they inflated within-trial precision by integer multiplication factors, drawing new effect estimates from the same random-effects model but with sampling variance divided by M. The between-study variance stayed fixed. Only the precision of each trial increased.

The results are stark. At an inflation factor of four, tau-squared was estimated at 0.008 — essentially the same, just random fluctuation — while I-squared rose to twenty-nine point two percent and Q climbed to ninety-eight with a p-value now crossing the significance threshold at 0.014. At a factor of sixteen, tau-squared was 0.027, still in the same range, but I-squared had reached eighty-four point eight percent and Q was four hundred fifty-four.

At a factor of sixty-four, tau-squared was 0.028 — indistinguishable from the original — and I-squared was ninety-six percent. Q had reached one thousand seven hundred eight. The statistic that was supposed to measure between-study differences was instead tracking how large the trials were.

The authors then checked whether this pattern holds across real meta-analyses, not just in a single simulated one. In a sample of one hundred fifty-seven binary-endpoint meta-analyses, they fitted a linear model predicting I-squared from two variables: tau-squared and the logarithm of the median study size. Both were strongly significant.

The coefficient for logarithm of median study size was eight point five — meaning each doubling of typical study size pushed I-squared up by about six points — and the model explained sixty-six percent of the variance in I-squared across those one hundred fifty-seven analyses. Tau-squared matters too, but the point is that study size matters independently, and the field has generally not accounted for that.

There is a further wrinkle Rücker and colleagues note. If you have a large trial whose estimated effect happens to be close to the pooled average, removing it is one of the most efficient ways to reduce I-squared. That should give pause.

A statistic that goes down when you exclude a well-powered, on-average trial is not straightforwardly measuring the heterogeneity of the evidence base.

So what should researchers use instead? Tau-squared. The between-study variance lives on the outcome scale — the same scale as the treatment effects themselves.

That means clinicians can ask a clinically meaningful question: is the spread of true effects across studies wide enough to change a treatment decision? Rücker and colleagues give a concrete anchor for odds-ratio outcomes. If the true effects across studies ranged from an odds ratio of 0.8 to 1.0 to 1.25, those values on the log scale correspond to tau — the standard deviation of true effects — of 0.22, which means tau-squared of roughly 0.05.

A reviewer can look at their estimated tau-squared, compare it to that threshold, and ask whether the spread matters clinically. Crucially, that comparison doesn't change just because the trials enrolled more patients.

The DerSimonian-Laird estimator is the standard tool for estimating tau-squared, and it is what Rücker and colleagues use throughout. They are candid that tau-squared has its own uncertainty — in the thrombolysis simulations, it fluctuated noticeably across runs, and the literature on heterogeneity estimation acknowledges that with few studies, any variance estimate will be imprecise. But imprecision is a reason to interpret carefully, not a reason to prefer a measure that is systematically biased by study size.

An uncertain answer to the right question beats a confident answer to the wrong one.

The implications reach into the infrastructure of evidence-based medicine. I-squared appears in almost every published systematic review. It is used by guideline committees, by journal reviewers deciding whether a meta-analysis is valid, and by clinicians trying to understand whether a pooled result applies to their patient.

If large, well-powered trials routinely push I-squared toward alarming values even when the actual between-study differences are modest, the field may be declining to pool evidence that should be pooled — or raising false alarms that send researchers hunting for moderators that don't exist.

The practical message from Rücker and colleagues is direct: when you read a meta-analysis, find the tau-squared, not the I-squared. Ask what the estimated spread of true effects is and whether that spread is large enough to change what you would do clinically. If tau-squared is small relative to a meaningful clinical difference, pooling is defensible — regardless of whether I-squared reads thirty percent or ninety percent.

The percentage that has come to dominate heterogeneity assessment is telling you something, but it is not telling you what most readers think it is.

This lecture was created by ennepō.

Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field.

Read when you can. Listen when you want to.

More in Decision Sciences