Publication Bias in Reports of Animal Stroke Studies Leads to Major Overstatement of Efficacy
Imagine reading a stack of papers about new stroke treatments in animals and thinking, wow, most of these look promising. Now imagine that a quiet stack of other experiments—equally real, just never published—tells a duller story. That gap, the difference between what gets printed and what gets left in a drawer, can bend the arc of a field.
That’s the heart of what Malcolm Macleod, Emily Sena, and colleagues set out to measure in animal models of stroke, not as a hunch, but with numbers.
They leaned on CAMARADES, a collaboration that, since 2004, has been curating systematic reviews of animal data. By August 2008, CAMARADES had details from 16 interventions tested in animal stroke models: 525 data sources drawn from 514 publications and 11 unpublished communications, describing 1,359 experiments in 19,956 animals. Here’s the first eyebrow-raiser.
Only 2 percent of publications reported no significant effect on infarct volume—the size of the brain area killed by the stroke—and just 1.2 percent failed to report at least one significant finding. In a universe where most experiments should miss as often as they hit, that skew is a red flag.
Sena and colleagues narrowed their lens to infarct-size outcomes for a reason. If a study measured multiple time points or reused the same cohorts in different arms, that could double-count signal. So they recorded the last outcome time point only, and if a cohort popped up more than once—say, in a combination-therapy paper—they censored repeats so each cohort appeared once in the pooled analysis.
Under the hood, they used a weighted, stratified mean-difference meta-analysis, storing, for every experiment, the effect size and its standard error. In other words, they built a dataset designed to say, as cleanly as possible, how much treatments seemed to shrink stroke damage in animals.
But the central question wasn’t just what the average effect was. It was how much that average was juiced by selective publication. So they reached for the standard toolkit.
Funnel plots first, which look at whether small, imprecise studies scatter equally above and below the mean, as they should if there’s no bias. Then Egger regression, which adds a line through that cloud; if the intercept is pulled away from zero, it signals an excess of imprecise studies with large effects. And, crucially, the Duval and Tweedie trim-and-fill method, implemented in the METATRIM module for STATA.
Trim-and-fill literally trims the studies that create asymmetry, estimates how many opposite, missing studies would balance the funnel, fills those in, and recalculates the overall effect.
What did the signals say? They shouted. Egger regression suggested publication bias in all 16 intervention datasets, with a positive intercept every time.
Trim-and-fill was a touch more conservative, suggesting bias for 10 of the 16 interventions. Across the entire pooled corpus, the method imputed 214 missing experiments—about one in seven beyond those identified—suggesting that a substantial slice of animal work never made it to print.
The cost of that missing slice shows up in the headline number. Unadjusted, the pooled efficacy across all interventions looked like a 31.3 percent reduction in infarct volume. After the trim-and-fill adjustment, that dropped to 23.8 percent.
That’s an absolute drop of about 7.5 percentage points and a relative overstatement of roughly one-third. Statistically, the difference isn’t a close call; the p-value sits well below the usual thresholds, below 0.0001. Practically, it means an intervention that looks like it shaves a third off stroke damage might actually shave closer to a quarter, once the invisible studies have their say.
It wasn’t just the pooled picture. On a per-intervention basis, bias left fingerprints too. In seven of the ten interventions where trim-and-fill detected asymmetry, the adjusted effect was significantly lower than the conventional meta-analysis.
The range was striking. Melatonin’s relative overstatement was modest, around 2.7 percent, while estrogens were at the other extreme, with more than a doubling—124 percent—relative overstatement. The absolute drops spanned roughly 1 to 15 percentage points depending on the intervention.
That spread matters. It says bias doesn’t just push in one direction; it pushes harder in some corners than others.
A couple of workhorse interventions also help anchor the scale. Hypothermia drew on 98 data sources, 222 experiments, and 3,256 animals, with an observed infarct-size reduction of around 43.5 percent, bracketed between roughly 40 percent and 47 percent. Tissue plasminogen activator—tPA—had an even larger footprint: 105 data sources, 256 experiments, and 4,029 animals, with an observed effect around 22.5 percent, bounded between about 19 percent and 26 percent.
These big datasets don’t just tell us two therapies look different; they also illustrate how publication bias got traction in places where there was a lot to sift through. For tPA, trim-and-fill suggested just 5 percent of experiments were "missing," while for interleukin-1 receptor antagonist, it was closer to 36 percent.
Now, you might be thinking, could funnel asymmetry come from something other than publication bias? Absolutely. As van der Worp and colleagues have emphasized, asymmetry can reflect differences in study quality, biology, or measurement noise.
Sena’s team leaned into that by scoring each study on a ten-item checklist—random allocation, blinding of ischemia induction and outcome assessment, temperature control, using an appropriate animal model, avoiding neuroprotective anesthetics, sample-size calculation, compliance with welfare regulations, conflict-of-interest statements, and whether it was peer-reviewed. Studies that checked more boxes tended, on average, to show smaller treatment effects. That’s consistent with a broader observation in preclinical work: better design can mean smaller, more reliable effects.
But there wasn’t a simple, straight-line link between a study’s statistical precision and its methodological quality, and no clear pattern that bigger literatures necessarily overstated more. In other words, multiple forces are shaping the funnel.
Still, the weight of the evidence stacks in one direction. Egger says there’s an excess of small, big-claim studies. Trim-and-fill imputes more than two hundred missing experiments.
And when you add those ghost studies back in, the pooled benefit shrinks by about a third. Add one more sobering number: if those unpublished experiments tracked the typical animal-per-experiment counts, you’re looking at roughly 3,600 animals whose data never entered the conversation. That’s not just a statistical concern. It’s an ethical one.
There’s also a design choice tucked into all of this that’s easy to miss but crucial. The analysis focused on infarct volume as the outcome because it’s one of the most common, quantifiable readouts in stroke models. Some interventions report behavior or survival, others multiple time points with magnetic resonance imaging, and those heterogeneities are fertile ground for duplicate counting and analytic drift.
By fixing on infarct size and standardizing how time points were selected and cohorts were handled, Sena and colleagues tried to take away those easy artifactual wins. It doesn’t magically harmonize small, heterogeneous studies, but it makes the pooled estimate less slippery.
Methodologically, this wasn’t a fishing expedition in obscure corners of the literature. The CAMARADES reviews that fed the analysis used broad searches, explicit inclusion and exclusion criteria, and dual screening. By 2008, they had captured 11 of the 14 meta-analyses of animal stroke studies published to that point.
The data management system recorded effect sizes and their uncertainty, and the repository remained open to public access on request. That doesn’t eliminate bias—if anything, it underscores why bias needs to be measured—but it means the pool being analyzed is as representative as the field had assembled at the time.
So what do we do with this? First, we recalibrate how we read preclinical stroke meta-analyses. If a pooled effect says thirty-one percent, the bias-adjusted reality might be closer to twenty-four.
That kind of translation matters when we’re deciding which therapies to push toward costly, risky human trials. Second, we start normalizing bias detection in preclinical synthesis. Funnel plots, Egger regression, trim-and-fill—none are perfect, but together they give us a cross-check.
Reporting the unadjusted and adjusted numbers side by side is a service, not an indictment.
Third—and this is the drum both Sena and van der Worp have been beating—we build the infrastructure to make nonpublication harder. A central register of animal experiments, grouped by topic, even anonymized where needed, would let the field see the denominator, not just the numerator. Think conference abstracts logged with a simple outcome summary.
Think preregistration, or at least registration, of animal studies so we know what was tried. Even if the results never turn into a paper, we’d have a record that can be counted in a synthesis.
There are limits worth holding in mind. Funnel asymmetry doesn’t wear a name tag that says "I am publication bias." Small studies aren’t just small; they’re often different in design, species, strain, or timing. And animal stroke models aren’t the whole preclinical world.
But as Sena and colleagues argue, there’s little reason to think stroke is unique here. If anything, it’s a test case for a broader, uncomfortable truth across preclinical science: we tend to see the exciting stuff, not the full story.
Let’s land it. In a large, carefully assembled dataset of animal stroke experiments, the published record looks brighter than the reality you get when you account for the studies that likely never saw the light of day. The best estimate is that around one in seven experiments went unpublished.
That inflation bumps the pooled efficacy from something like a quarter to something like a third. And because those estimates shape what we fund, what we test in people, and how we treat animals in research, getting them right isn’t a pedantic concern. It’s the difference between being guided by a map and by a mirage.
As a field, we can do better. Build the registries. Expect bias assessments in preclinical meta-analyses.
Treat small, underpowered literatures with special care. And when a paper reports a big effect, ask not just how big it is, but how many quiet, missing studies it might be standing on.
Related lectures
- Brucella abortus Uses a Stealthy Strategy to Avoid Activation of the Innate Immune System during the Onset of Infection
- Are animal models predictive for humans?
- Global Estimate of Human Brucellosis Incidence
- Evaluation of EMLA Cream for Preventing Pain during Tattooing of Rabbits: Changes in Physiological, Behavioural and Facial Expression Responses