Publication bias and the limited strength model of self-controlhas the evidence for ego depletion been overestimated?
In 1998, Roy Baumeister's lab published a finding so intuitive that it felt like it had to be true. Willpower, as argued in the paper, works like a muscle. Use it on one task, and you'll have less of it for the next. That single idea spawned over two hundred studies, a major meta-analysis, and a theory that spread from academic psychology into nutrition advice and self-help culture. Then, in 2014, Evan Carter and Michael McCullough looked at all that evidence together. What they found wasn't a muscle running out. It was a literature running on empty. The theory itself is called the limited strength model of self-control. Its core claim is clear: acts of self-control draw on a finite resource, and depleting that resource on an initial task leaves less of it available for subsequent tasks. Researchers referred to this as the depletion effect or ego depletion. The sequential-task paradigm translated it into dozens of doable experiments. They would force someone to override a habit or resist a temptation on task one, and then measure how they perform on an unrelated self-control task afterward. If performance dropped, that was evidence of depletion. The paradigm worked across a surprising range of behaviors: impulse control, cognitive effort, active decision-making, social regulation.
In 2010, Martin Hagger and colleagues synthesized nearly two hundred of these experiments into a meta-analysis — a statistical pooling of results — and the headline number was a d of 0.62. In psychology, a d of 0.62 is a medium effect, nothing to dismiss. The confidence interval was tight: from 0.57 to 0.67. The conclusion was that ego depletion was real, consistent, and worth building theories on. Labs ran with it. The model accumulated elaborations, moderators, mediators, and applied extensions. The foundation seemed solid. The problem was hiding inside the meta-analysis itself. Publication bias is the tendency for studies with positive, significant results to be published at higher rates than studies with null or negative results. Small-study effects is the broader term for when smaller studies yield systematically larger effect estimates than bigger ones. Publication bias is one common cause, but not the only one. Either way, if a meta-analysis pools a literature shaped by these forces, the resulting estimate can be inflated far above the true effect, regardless of how many studies went into it. A funnel plot makes this visible. Each study sits at a point defined by its effect size and its standard error, which is a measure of precision. Big precise studies cluster tightly near the true effect at the top of the plot, while smaller, noisier studies fan out below.
If the literature is unbiased, that fan is symmetric. If the fan is lopsided, with small studies clustering heavily on the side showing large positive effects and the opposite side conspicuously empty, that asymmetry tells you something is missing. A variant called a contour-enhanced funnel plot shades the region where results would be statistically non-significant, making it easy to see if the asymmetry falls precisely where null results would appear. When it does, the file-drawer problem — unpublished null results sitting on hard drives — is the most plausible culprit. Carter and McCullough went back to Hagger and colleagues' dataset and applied four corrective tools. The first was the binomial excess-significance test, which asks whether the number of statistically significant results in a literature exceeds what you'd expect given the average statistical power of the studies. Across every subsample in this literature, average power ranged from only 0.31 to 0.69, well below the conventional 0.80 benchmark. In all but one subsample, the observed number of significant results exceeded expectations. The second tool was trim-and-fill, a method that estimates how many studies are missing from the funnel plot, imputes them, and recalculates the overall effect. For the full sample, trim-and-fill needed to impute seventy-three missing studies to restore symmetry. After adding them, the fixed-effect estimate dropped to 0.48 and the random-effects estimate to 0.50.
That represents a meaningful reduction, but trim-and-fill has known limitations, particularly when heterogeneity between studies is high. The more powerful tools were regression-based. The Precision-Effect Test, or PET, models the relationship between each study's effect size and its standard error. The intercept of that regression, called b-zero, represents the effect size extrapolated to a perfectly precise study with a standard error of zero. PET is most reliable when the true effect is actually zero. The Precision-Effect Estimate with Standard Error, or PEESE, uses variance rather than standard error as its predictor and performs better when the true effect is genuinely non-zero. The recommended conditional strategy, called PET-PEESE, is to run PET first. If its intercept is statistically distinguishable from zero, then you can trust PEESE's estimate. If not, trust PET's. Here is where the result lands. When Carter and McCullough applied PET to the full depletion dataset, the intercept b-zero was negative 0.10, with a ninety-five percent confidence interval from negative 0.23 to 0.02. That interval crosses zero.
The effect, after correction, is indistinguishable from zero. Because PET's intercept was not significant, the PET-PEESE rule says to stop there and treat PET's estimate as the least-biased available number. The slope coefficient testing funnel asymmetry was highly significant, with a p-value below 0.001, meaning the correlation between effect size and study precision was real and strong. The contour-enhanced funnel plots confirmed the picture: asymmetry was concentrated precisely in the region where non-significant results would fall. Sit with that for a second. The headline meta-analytic estimate was a d of 0.62. The bias-corrected estimate is not different from zero. That is not a marginal trim; that is the entire effect disappearing. To understand how that happens, Carter and McCullough cite simulation work by Bakker and colleagues showing that publication bias alone, without any deliberate misconduct, can inflate a truly null literature to an observed meta-analytic effect of d of 0.35. Add what researchers call degrees of freedom — flexible decisions about when to stop collecting data, which outcomes to report, and which participants to exclude — and the simulated inflated estimate rises to 0.48, with highly significant funnel asymmetry. John and colleagues estimated that seventy-two percent of psychology researchers have used optional stopping, seventy-eight percent have failed to report all dependent measures, and sixty-two percent have excluded data post-hoc.
These aren't rare bad actors. These are widespread practices that individually seem minor and collectively can manufacture a literature. The methodological lesson that Carter and McCullough draw is specific and actionable, not a counsel of despair. The correction tools they applied — PET, PEESE, and the conditional PET-PEESE — can detect and partially correct for small-study effects, and they often outperform older methods like trim-and-fill. But they work less reliably when between-study heterogeneity is large, and they are not a substitute for producing unbiased evidence in the first place. For that, the authors call for coordinated, large, pre-registered direct replications with committed data-sharing. The scale of what "large" means here is striking. If the true depletion effect were as small as a d of 0.25 — the PEESE intercept for the full sample — you would need roughly two hundred fifty-two participants per condition to achieve eighty percent statistical power. The average study in Hagger and colleagues' meta-analysis had about twenty-seven participants per condition. The entire edifice of ego-depletion evidence was built on studies roughly one-tenth the size needed to reliably detect even a modest effect. Carter and McCullough are careful not to overclaim their conclusion. They do not say the depletion effect is definitively zero. What they say is that the existing evidence is not convincing.
After applying the best available bias-correction methods, the effect is not distinguishable from zero, and the field should treat this as unresolved rather than established. That distinction matters because it dictates what should happen next. The priority, they argue, should not be elaborating the limited strength model with new moderators or neuroscientific mechanisms. It should be answering the foundational question first: does the basic effect exist at all? That is a harder ask than it sounds. Ego depletion became embedded in popular psychology, in recommendations about decision fatigue, and in narratives about why judges rule more harshly before lunch. None of that is invalidated by one paper. But what Carter and McCullough showed is that the scientific scaffolding underneath those applications is not load-bearing — not yet. The effect might be real. It might be small. It might be zero. Two hundred published studies, pooled uncritically, cannot tell us which. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Collective Trauma and the Social Construction of Meaning
- Psychological Impact and Associated Factors During the Initial Stage of the Coronavirus (COVID-19) Pandemic Among the General Population in Spain
- Psychoeducation for depression, anxiety and psychological distress: a meta-analysis
- The impact of the COVID-19 epidemic on mental health of undergraduate students in New Jersey, cross-sectional study
- Embodied Cognition is Not What you Think it is
- Quality of Acute Psychedelic Experience Predicts Therapeutic Efficacy of Psilocybin for Treatment-Resistant Depression