The heterogeneity statistic I2 can be biased in small meta-analyses

Paul T. von HippelView original
OverviewBalancedbennett voice
You're reading a meta-analysis. Seven studies, a reported I-squared of 21 percent. The authors describe this as "low to moderate" heterogeneity and proceed accordingly. Should you trust that number? Paul von Hippel asked exactly that question, and the math he found in answer should make you read the next meta-analysis you encounter with considerably more skepticism. Meta-analysis is the practice of pooling results from multiple studies into a single summary estimate. When different studies come up with different answers — which they almost always do — that variation has two possible sources. Some of it is just random sampling noise; different studies pull from different slices of the population by chance. But some of it reflects real differences between the studies themselves: different patient populations, different doses, and different ways of measuring the outcome. That second kind is called heterogeneity, and it matters enormously for interpretation. A combined estimate from homogeneous studies can plausibly generalize to new settings. A combined estimate from heterogeneous studies might be hiding important variation — an average that nobody in the real world actually experiences. So meta-analysts quantify heterogeneity. The classical tool is Cochran's Q, but Q has a fundamental problem: its value depends heavily on how many studies you have. With many studies, Q will flag heterogeneity as statistically significant even when the underlying differences are trivial. With few studies, it lacks the statistical power to detect real heterogeneity. To get around this, Higgins and Thompson introduced I-squared — the fraction of the total variance in study effect estimates that is due to real between-study differences rather than sampling noise. I-squared runs from zero to one hundred percent. It's more interpretable than Q, and it was designed to be less sensitive to the number of studies. In practice, it has become the standard way to communicate heterogeneity in published meta-analyses. Here's the problem. Small meta-analyses are not exotic. They are the norm. Von Hippel reports that in the Cochrane Library — the gold standard repository of medical evidence synthesis — the median number of studies per meta-analysis is seven or fewer. Some summaries of Cochrane reviews have found medians as low as three. And the median reported I-squared in that same library is 21 percent. So the most commonly used heterogeneity statistic is being most commonly applied in the exact conditions where, it turns out, it is most likely to mislead. Von Hippel computed the bias of I-squared analytically, using Mathematica to calculate the mathematical expectation of the statistic under known distributional assumptions about Cochran's Q. This lets him compare what I-squared is expected to report against what the true heterogeneity actually is. The result is a bias profile, and it cuts in two opposite directions depending on whether true heterogeneity is small or large. When true heterogeneity is small, I-squared tends to overestimate it. With seven studies and no real heterogeneity at all, the expected I-squared is about 12 percent — an overstatement of 12 percentage points right out of the gate. When true heterogeneity is large, I-squared tends to underestimate it. With seven studies and a true heterogeneity of 80 percent, the expected I-squared is only around 52 percent — an understatement of 28 percentage points. The statistic doesn't consistently exaggerate or consistently minimize. It compresses both extremes toward the middle. Small truths get inflated; large truths get attenuated. The bias changes sign somewhere around 20 percent, which is almost exactly where the median Cochrane estimate sits. That's not a coincidence that should comfort you — it means the published literature is clustered right at the ambiguous zone. Understanding why this happens requires a short look at how I-squared is actually computed. The naïve estimator — the formula Higgins and Thompson originally proposed — is essentially one minus the ratio of Q's degrees of freedom to Q itself. In words, you take the observed Q statistic, see how much it exceeds what you'd expect by chance, and express that as a proportion. The trouble is this formula can produce negative values whenever Q falls short of its expected value, which happens more than half the time when the number of studies is small. Since a fraction of variance can't be negative, researchers truncate: if the formula gives you a negative number, you report zero instead. That truncation sounds like a fix, but it introduces its own bias. Every time the true value is close to zero and sampling variation pushes the naïve estimate below zero, you round up to zero. Over many hypothetical repetitions of the same meta-analysis, those rounded-up zeros inflate the average. The bias near zero is partly the cost of enforcing a sensible floor. Von Hippel also draws a careful distinction between the estimand and the estimator — between what you're trying to measure and the formula you use to measure it. The true quantity is denoted I-squared: between-study variance divided by the sum of between-study variance and average within-study variance. I-squared is trying to estimate that. But because of the truncation problem and the behavior of Q in small samples, I-squared and I-squared diverge systematically. The model framing matters too. Under a random-effects model — where studies are assumed to be drawn from a broader population of possible studies — the bias at high true heterogeneity is severe. Under a fixed-effects model — where you treat the studies in hand as the entire population of interest — the same scenario produces an essentially unbiased estimate. The choice of model is not just philosophical. It changes the bias profile of the very statistic you're reporting. So what should you do? Von Hippel's recommendation is direct: when a meta-analysis has few studies, report a confidence interval around I-squared rather than relying on the point estimate alone. A confidence interval makes the uncertainty visible. In Cochrane meta-analyses, a typical ninety-five percent confidence interval around I-squared runs from roughly zero to 60 percent. That means a reported value of 21 percent is statistically consistent with anywhere from essentially no heterogeneity to heterogeneity accounting for the majority of between-study variance. Those are very different scientific situations. A point estimate of 21 percent, standing alone, conceals that entire range. Von Hippel also highlights a parallel issue with tau-squared — the raw between-study variance that I-squared is built from. Any estimator of tau-squared constrained to be nonnegative will share the same upward bias near zero. The problem isn't specific to the I-squared rescaling. It runs deeper, into the architecture of how heterogeneity variance is estimated from small samples. There's a practical tension here that von Hippel acknowledges directly. Confidence intervals for I-squared are not routinely reported. He cites recent meta-analyses in journals including Epidemiology and the American Journal of Epidemiology, as well as the Cochrane Library itself, where point estimates appear without intervals. The methods for computing those intervals exist and perform well — they just haven't become standard practice. What this means for a listener who reads research — or relies on guidelines built from it — is concrete. When you see a meta-analysis with fewer than ten studies, the I-squared is telling you something, but not as precisely or as accurately as it appears. An I-squared of 50 percent with seven studies might reflect true heterogeneity of 20 percent or 80 percent. An I-squared of zero — which appears in roughly one quarter of published meta-analyses — might reflect genuine homogeneity or might simply reflect a noisy estimate that got rounded to the floor. The point estimate alone cannot tell you which. Von Hippel's case is not that I-squared should be abandoned. It's that it should be interpreted with the same honest accounting of uncertainty that we demand from other statistical estimates. Means come with standard errors. Regression coefficients come with confidence intervals. Heterogeneity statistics should too. The median meta-analysis in the Cochrane Library has seven studies. The median reported I-squared is 21 percent. The typical confidence interval around that estimate spans from zero to 60 percent. That interval is the truth. The point estimate is a guess — and a biased one. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

You're reading a meta-analysis. Seven studies, a reported I-squared of 21 percent. The authors describe this as "low to moderate" heterogeneity and proceed accordingly. Should you trust that number? Paul von Hippel asked exactly that question, and the math he found in answer should make you read the next meta-analysis you encounter with considerably more skepticism. Meta-analysis is the practice of pooling results from multiple studies into a single summary estimate. When different studies come up with different answers — which they almost always do — that variation has two possible sources. Some of it is just random sampling noise; different studies pull from different slices of the population by chance. But some of it reflects real differences between the studies themselves: different patient populations, different doses, and different ways of measuring the outcome. That second kind is called heterogeneity, and it matters enormously for interpretation. A combined estimate from homogeneous studies can plausibly generalize to new settings. A combined estimate from heterogeneous studies might be hiding important variation — an average that nobody in the real world actually experiences. So meta-analysts quantify heterogeneity. The classical tool is Cochran's Q, but Q has a fundamental problem: its value depends heavily on how many studies you have. With many studies, Q will flag heterogeneity as statistically significant even when the underlying differences are trivial.

With few studies, it lacks the statistical power to detect real heterogeneity. To get around this, Higgins and Thompson introduced I-squared — the fraction of the total variance in study effect estimates that is due to real between-study differences rather than sampling noise. I-squared runs from zero to one hundred percent. It's more interpretable than Q, and it was designed to be less sensitive to the number of studies. In practice, it has become the standard way to communicate heterogeneity in published meta-analyses. Here's the problem. Small meta-analyses are not exotic. They are the norm. Von Hippel reports that in the Cochrane Library — the gold standard repository of medical evidence synthesis — the median number of studies per meta-analysis is seven or fewer. Some summaries of Cochrane reviews have found medians as low as three. And the median reported I-squared in that same library is 21 percent. So the most commonly used heterogeneity statistic is being most commonly applied in the exact conditions where, it turns out, it is most likely to mislead.

Von Hippel computed the bias of I-squared analytically, using Mathematica to calculate the mathematical expectation of the statistic under known distributional assumptions about Cochran's Q. This lets him compare what I-squared is expected to report against what the true heterogeneity actually is. The result is a bias profile, and it cuts in two opposite directions depending on whether true heterogeneity is small or large. When true heterogeneity is small, I-squared tends to overestimate it. With seven studies and no real heterogeneity at all, the expected I-squared is about 12 percent — an overstatement of 12 percentage points right out of the gate. When true heterogeneity is large, I-squared tends to underestimate it. With seven studies and a true heterogeneity of 80 percent, the expected I-squared is only around 52 percent — an understatement of 28 percentage points. The statistic doesn't consistently exaggerate or consistently minimize. It compresses both extremes toward the middle. Small truths get inflated; large truths get attenuated. The bias changes sign somewhere around 20 percent, which is almost exactly where the median Cochrane estimate sits. That's not a coincidence that should comfort you — it means the published literature is clustered right at the ambiguous zone.

Understanding why this happens requires a short look at how I-squared is actually computed. The naïve estimator — the formula Higgins and Thompson originally proposed — is essentially one minus the ratio of Q's degrees of freedom to Q itself. In words, you take the observed Q statistic, see how much it exceeds what you'd expect by chance, and express that as a proportion. The trouble is this formula can produce negative values whenever Q falls short of its expected value, which happens more than half the time when the number of studies is small. Since a fraction of variance can't be negative, researchers truncate: if the formula gives you a negative number, you report zero instead. That truncation sounds like a fix, but it introduces its own bias. Every time the true value is close to zero and sampling variation pushes the naïve estimate below zero, you round up to zero. Over many hypothetical repetitions of the same meta-analysis, those rounded-up zeros inflate the average. The bias near zero is partly the cost of enforcing a sensible floor. Von Hippel also draws a careful distinction between the estimand and the estimator — between what you're trying to measure and the formula you use to measure it. The true quantity is denoted I-squared: between-study variance divided by the sum of between-study variance and average within-study variance. I-squared is trying to estimate that.

But because of the truncation problem and the behavior of Q in small samples, I-squared and I-squared diverge systematically. The model framing matters too. Under a random-effects model — where studies are assumed to be drawn from a broader population of possible studies — the bias at high true heterogeneity is severe. Under a fixed-effects model — where you treat the studies in hand as the entire population of interest — the same scenario produces an essentially unbiased estimate. The choice of model is not just philosophical. It changes the bias profile of the very statistic you're reporting. So what should you do? Von Hippel's recommendation is direct: when a meta-analysis has few studies, report a confidence interval around I-squared rather than relying on the point estimate alone. A confidence interval makes the uncertainty visible. In Cochrane meta-analyses, a typical ninety-five percent confidence interval around I-squared runs from roughly zero to 60 percent. That means a reported value of 21 percent is statistically consistent with anywhere from essentially no heterogeneity to heterogeneity accounting for the majority of between-study variance. Those are very different scientific situations. A point estimate of 21 percent, standing alone, conceals that entire range.

Von Hippel also highlights a parallel issue with tau-squared — the raw between-study variance that I-squared is built from. Any estimator of tau-squared constrained to be nonnegative will share the same upward bias near zero. The problem isn't specific to the I-squared rescaling. It runs deeper, into the architecture of how heterogeneity variance is estimated from small samples. There's a practical tension here that von Hippel acknowledges directly. Confidence intervals for I-squared are not routinely reported. He cites recent meta-analyses in journals including Epidemiology and the American Journal of Epidemiology, as well as the Cochrane Library itself, where point estimates appear without intervals. The methods for computing those intervals exist and perform well — they just haven't become standard practice. What this means for a listener who reads research — or relies on guidelines built from it — is concrete. When you see a meta-analysis with fewer than ten studies, the I-squared is telling you something, but not as precisely or as accurately as it appears. An I-squared of 50 percent with seven studies might reflect true heterogeneity of 20 percent or 80 percent. An I-squared of zero — which appears in roughly one quarter of published meta-analyses — might reflect genuine homogeneity or might simply reflect a noisy estimate that got rounded to the floor. The point estimate alone cannot tell you which.

Von Hippel's case is not that I-squared should be abandoned. It's that it should be interpreted with the same honest accounting of uncertainty that we demand from other statistical estimates. Means come with standard errors. Regression coefficients come with confidence intervals. Heterogeneity statistics should too. The median meta-analysis in the Cochrane Library has seven studies. The median reported I-squared is 21 percent. The typical confidence interval around that estimate spans from zero to 60 percent. That interval is the truth. The point estimate is a guess — and a biased one. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Decision Sciences