The Meaningfulness of Effect Sizes in Psychological ResearchDifferences Between Sub-Disciplines and the Impact of Potential Biases
Effect sizes are the exchange rate of psychology. They tell you not just whether something happens but how much it happens. For decades, we’ve relied on two ways to make sense of them.
First, compare your effect to what the field has seen before. Second, use Cohen’s global yardsticks—small, medium, large. Simple, tidy, and reassuring. But, as it turns out, also shaky.
Here’s the problem in plain terms. Both routes assume the literature is a fair mirror, that what’s been published presents a decent picture of the true effects out there. However, publication bias, flexible analyses, and selective reporting distort that mirror.
Additionally, psychology isn’t one thing. A memory study with the same correlation value as a social-priming study might mean something very different because the designs and the noise are different. So, if the mirror is warped and the yardstick is one-size-fits-all, how do we interpret what we see?
Thomas Schäfer and Marcus Schwarz set out to test those two pillars directly. First, are published effects in psychology larger than the true population effects because of bias? Second, even if you had unbiased estimates, could global benchmarks possibly make sense across wildly different sub-fields and designs?
They didn’t argue this in the abstract; they built a sampling plan to represent psychology as it is actually practiced and asked the data to speak.
The backbone of their approach was breadth and randomness. They mapped psychology using nine Social Sciences Citation Index sub-disciplines—applied, biological, clinical, developmental, educational, experimental, multidisciplinary, psychoanalysis, and social. Then they rolled the dice.
For each area, they randomly selected 10 journals, pulled all volumes and issues, and then randomly picked 10 articles from each journal. That’s 900 empirical effects, a century’s worth of papers up through 2018, sampled without cherry-picking. None of those articles were preregistered, by the way.
To compare contemporary practices with a bias-resistant workflow, they then gathered every preregistered empirical article they could find at that time. The tally came to 93. Roughly half were replications, including several multilab efforts, while the rest were original studies.
Sixteen were registered reports, where the editorial decision is made before results exist. All of this sat on an open repository, so anyone could rerun the analysis or slice it differently.
Now, effects in psychology wear a lot of clothes—Cohen’s d here, partial eta-squared there, t's and F's everywhere. So, Schäfer and Schwarz stripped them down to a common language: Pearson’s r. They converted Cohen’s d, Hedges' g, and partial eta-squared to r, and when needed, derived r from t and F statistics.
Crucially, they split the world by design because within-subjects experiments—where the same people provide multiple measurements—cancel a lot of noise that between-subjects comparisons can’t touch. For patterns, they used a very smooth loess fit—think of it as a flexible trend line—with a high smoothing setting, and they wrapped their medians and other summaries in bootstrap confidence intervals to gauge uncertainty.
So, what did the mirror show when you polish it? The headline is stark. In the broad, non-preregistered literature, the median effect size was a correlation value of zero point thirty-six.
In the preregistered set, the median was a correlation value of zero point sixteen. That’s about a twofold drop. If you’ve been worried that biases inflate what gets published, this is what that inflation looks like when you measure it.
And it wasn’t just the size of effects. The share of statistically significant results was lower in the preregistered studies as well, exactly what you’d expect when you stop chasing p-values.
Design mattered a lot. Within the preregistered studies, within-subjects designs produced much larger effects than between-subjects ones—medians of zero point thirty-one versus zero point twelve—and that gap wasn’t a fluke. The difference was statistically reliable, with a t-value a little under four and a p-value around three thousandths.
You saw the same pattern in the non-preregistered set too. That tracks with intuition: when each participant is their own control, you strip out between-person variability, and the signal stands taller above the noise.
Then there’s the question of whether psychology is one landscape or many. When Schäfer and Schwarz laid out effect sizes by sub-discipline, the ranges barely recognized each other. Experimental and biological psychology were on the high side, while social and developmental psychology were lower.
In fact, the typical ranges for social and biological psychology didn’t even overlap. That alone is a death knell for global "small, medium, large" cutoffs. A medium effect in one neighborhood could be a giant in another.
Another pattern cuts across all of this: the relationship between sample size and the size of reported effects. In the broad, conventional literature, larger samples tended to come with smaller observed effects. That’s the footprint of publication bias—small studies with big results make it to print, while small studies with small results often don’t.
Even in the preregistered set, where that bias should weaken, the negative trend didn’t vanish entirely. The smooth trend line they fit captured that gentle slide. One thing that didn’t seem to matter was time.
Looking across the century, effect sizes in the non-preregistered literature didn’t drift up or down in any systematic way. Preregistered papers tended to recruit more people. The typical sample size in the broad set was under one hundred; in the preregistered studies, the median was around 267. Bigger samples led to smaller but sturdier effects.
Now, zoom out for a second. If you’ve been following the replication literature, none of this will sound alien. The Open Science Collaboration—hundreds of researchers re-running classic studies—reported that effects in direct replications were roughly half the size of the originals.
In their findings, the average correlation dropped from about zero point four to zero point two, and an average standardized mean difference, Cohen’s d, shrank from roughly zero point six to zero point fifteen. Schäfer and Schwarz are observing that same halving when they compare conventional publications to preregistration. Two lines of evidence, one story.
What does that mean for how we plan and interpret studies? It means that those global Cohen benchmarks are a broken compass right now. If you power your study assuming a medium effect because that’s what the literature seemed to show, and the true effect in your sub-field is a third of that, you will come up short—underpowered, under-sampled, and overconfident.
It also means that comparisons to typical effects drawn from the published record can mislead theory tests. If published effects are systematically inflated, a new finding that looks small by comparison might actually be right on target.
Let’s talk caveats, because they matter. The preregistered sample here is modest. Ninety-three studies is a start, not an endpoint, and preregistration itself comes in flavors—registered reports are more insulated from bias than a garden-variety preregistration that still undergoes conventional review.
There’s also the possibility of self-selection: some teams may be more likely to preregister certain kinds of questions. And the way multi-study papers were handled—taking the first main effect to avoid dependency—could tilt things a bit. None of that erases the inflation pattern, but it does put guardrails around how far we generalize it.
So, how do we do better? Schäfer and Schwarz argue for two shifts. First, stop pretending that one-size-fits-all benchmarks make sense across the whole field.
If you want a benchmark, build it within a homogeneous slice: your sub-discipline, your design, your measurement tradition. Second, wherever possible, report unstandardized effects alongside the standardized ones. Tell us what a one-unit change means in the real world of your construct.
It’s harder to compare across scales, but it grounds interpretation in psychology rather than in z-scores.
Power analysis follows the same logic. Use design-aware, domain-specific estimates, and be conservative. If preregistered work points to smaller true effects, plan for that reality rather than for the rosy picture painted by selective publication.
In practice, that means larger samples, especially for between-subjects designs where noise is your enemy. And then, of course, preregister. Registered reports are even better.
When the decision to publish happens before the results exist, the incentive to massage them evaporates, and our picture of the population effect gets cleaner.
There’s one more thread worth pulling. Schäfer and Schwarz didn’t find that effect sizes have been drifting over time in the conventional literature. That’s a reminder that the problem isn’t a sudden crisis; it’s a long, steady accumulation of bias.
The fix won’t be a single paper either. It’ll be a pipeline change—more preregistration, more registered reports, and a habit of reporting effects in ways that respect design differences and real-world meaning.
If you’re a listener who conducts studies, here’s the simple takeaway. Calibrate your expectations within your niche, not against a field-wide rule of thumb. If the best preregistered evidence in your corner of psychology points to a correlation value around zero point fifteen, don’t power for zero point thirty just because it feels nicer.
If you work within subjects, you can often get away with fewer participants than in between-subjects work, but don’t confuse that design advantage with a license to downsize indiscriminately. And if you’re synthesizing evidence, treat cross-sub-discipline comparisons of effect sizes the way you’d treat cross-species comparisons of metabolism: with caution and context.
The hopeful part is this: the solution scales. As more preregistered studies accumulate, we’ll be able to chart effect-size distributions that are both less biased and more local—by topic, by design, by measure. Those distributions can become the new benchmarks. Not universal, but useful. Not simple, but honest.
And that’s the pivot we’re experiencing. From a field where effect sizes were judged by a convenient global yardstick to one that interprets them in context, powered by transparent methods. Cohen gave us a language. Replication and preregistration are teaching us how to speak it precisely.
Related lectures
- Insomnia and the risk of depression: a meta-analysis of prospective cohort studies
- Mirror-Induced Behavior in the Magpie (Pica pica): Evidence of Self-Recognition
- The cross-national epidemiology of social anxiety disorder: Data from the World Mental Health Survey Initiative
- The Natural Statistics of Audiovisual Speech
- The Small World of Psychopathology
- Health-related quality of life in parents of school-age children with Asperger syndrome or high-functioning autism