The Extent and Consequences of P-Hacking in Science

Megan L. Head, Luke Holman, Robert Lanfear, Andrew T. Kahn, Michael D. JennionsView original
OverviewBalancedalloy voice
If you've ever wondered why the number 0.05 shows up like a speed limit sign in so many papers, here's the story. For almost a century, scientists have leaned on a ritual called null hypothesis significance testing. You pose a null hypothesis—no effect, no relationship—run your study, and compute a p-value: the probability, if the null were true, of seeing data as extreme as yours. Then comes the cliff. If the p-value falls below 0.05, the result gets stamped "significant" and is treated as a win. If it's above 0.05? It's often ignored. That cutoff was meant to separate signal from noise. In practice, it shapes careers, headlines, and what even counts as knowledge. When a single number carries that much weight, it starts to bend behavior. Head and colleagues gave a crisp name to the result: inflation bias, better known as p-hacking. You try different analysis choices—stop data collection when the graph looks good, test many outcomes and report the ones that pop, decide on outliers after a peek—and you keep the paths that land you under 0.05. Not necessarily with malice. The incentive structure does the nudging. Journals favor "significant" results. Funders and hiring committees count them. And there's a common confusion in the background: a small p-value can come from a tiny effect in a huge sample. It is not a measure of importance. All of this means that the pressure to deliver a p-value less than 0.05 is both cultural and structural. So how do you tell when a literature is signaling a real effect versus telling you what it thinks you want to hear? One surprisingly simple diagnostic is the p-curve, the distribution of reported significant p-values. Picture it this way. If there's a true effect out there, very small p-values accumulate more than modestly small ones, so the curve leans to the right—more values of 0.01 than near 0.05. If there's no effect at all and researchers are just publishing the "wins," the p-values among those wins tend to look flatter. Now add p-hacking. When researchers nudge near-misses over the line, you get a bump—an overabundance of p-values just below 0.05. That little cliff shapes the landscape. Crucially, the exact shape depends on two forces: whether a true effect exists and how hard people are pushing their analyses toward significance. Head, Holman, Lanfear, Kahn, and Jennions asked a big, blunt question: across science, do we see the signal of real effects, the fingerprints of p-hacking, or both? And do the conclusions we draw from meta-analyses—our best attempt at synthesis—survive those distortions? They built a scalable toolkit to find out. First, they text-mined p-values from open access papers across fields, separating the numbers reported in results sections from those in abstracts. Second, they reanalyzed the p-values that underlie published meta-analyses, where the questions are clearly defined and the data have already been assembled. The core test is elegant. For evidential value, they compare how many p-values fall between zero and 0.025 to how many land between 0.025 and 0.05. More in the lower bin means a right-skewed curve and a likely real effect. For p-hacking, they zoom in on the danger zone with a finer split: values of 0.04 to 0.045 versus values of 0.045 to 0.05. If the upper sliver is stuffed, that's a red flag. Statistically, they frame these as binomial questions: is the proportion in the upper bin different from 0.5? They use a binomial generalized linear model to test that intercept. They also sanity-check with a sign-test variant popularized by Uri Simonsohn and colleagues. To keep the text-mined data from being dominated by prolific writers, they bootstrap one p-value per paper section a thousand times. And when they quote proportions, the confidence intervals come from the Clopper-Pearson method, the exact one you get from binom.test in R. Across the text-mined literature, the first message is reassuring. Pulling p-values from results sections across 14 disciplines, the proportion in the upper evidential bin—those between 0.025 and 0.05—lands at 0.257. Abstracts tell a very similar story across 10 disciplines with 0.262. That's a right-skew, just as you'd expect if many reported effects are real. You can think of it like this: there are more very small p-values than near-threshold ones, which is hard to fake consistently without a true signal underneath. But the second message is sobering. Zooming in on the sliver just below 0.05, the proportion tips the other way: 0.546 in results and 0.537 in abstracts. That is, there are more p-values in the 0.045 to 0.05 window than in the 0.04 to 0.045. Exactly where you'd expect to see the fruits of garden-of-forking-paths decisions, you do. Many individual disciplines show this bump clearly. And here's the twist: when the team stripped out misreported p-values—those labeled as "p less than 0.05" that, on inspection, weren't—the p-hacking signal shrank dramatically in the discipline-wide summary. Misreporting itself is, of course, part of the problem, but it also exaggerates how strong the p-hacking bump looks when you scan the literature at scale. The second line of evidence focuses on the spine of our evidence base: meta-analyses. Here, Head and colleagues reassembled the p-curves from the primary data used in 12 published syntheses, each addressing a specific, pre-stated question. Again, you get that reassuring rightward lean. The pooled upper-bin proportion is 0.202 across these datasets, and nine of the twelve individual curves show evidential value on their own. The three holdouts tended to have small sample sizes, which likely just means they lacked the statistical power to show the skew, not that the underlying effects were absent. What about p-hacking within these meta-analytic datasets? Misreported p-values did crop up in the pile. When the team left those in, the fine-grained test near the threshold flagged a problem: the pooled proportion in the p equal to 0.045 to 0.05 window rose to 0.615 with a statistically significant test result. Remove the misreports, and that proportion drops to 0.489 and the signal evaporates statistically. One large meta-analysis by Jiang and collaborators provides a vivid example of this contrast; the apparent p-hacking signal hinged on the misreports. The takeaway is not that p-hacking disappears inside meta-analyses. It's that a chunk of what you detect can be an artifact of p-values that were mislabeled as "significant" in the first place. Let's pause and translate that into plain consequences. Across fields, there is real signal—more tiny p-values than near-threshold ones—so the scientific literature isn't just smoke and mirrors. At the same time, the neighborhood right under 0.05 is crowded, which is exactly where p-hacking would park its cars. And when you pool studies in a meta-analysis, that crowd can inflate the average effect size a bit. But because meta-analyses weight studies, the noisiest, smallest sample ones—the ones most vulnerable to p-hacking—tend to count for less. That doesn't make them immune. It does mean synthesis can, under decent conditions, dampen the worst inflations. Now, methods have limits, and Head and colleagues are candid about them. The two-bin evidential test is tuned to catch strong right-skews; it's less sensitive to subtle p-hacking. The hyper-focused 0.04 versus 0.045 split is great for zeroing in near the threshold, but making bins smaller boosts sensitivity at the cost of statistical power. And text-mined p-values mash together many different questions within a discipline. A right-skew there tells you that, in aggregate, the field is finding real effects. It can't guarantee that any single, well-defined hypothesis is on solid footing. Signals also depend on average effect sizes and study power. Underpowered domains can look less evidential even when the effects are real. With those caveats, the toolkit they showcase—p-curves, binomial modeling of adjacent bins, bootstrap sampling to avoid overcounting prolific authors, and exact binomial intervals—does something valuable. It turns a vague worry into testable shape. If you see an overall right-skew and a bump just below 0.05, you can quantify both: how much real signal there is, and how much near-threshold inflation you're likely facing. And because you can apply the same diagnostics to the primary data behind a meta-analysis, you can check whether synthesis itself is amplifying, damping, or surviving those pressures. What should we do about the pressures that generate the bump in the first place? No single reform will erase p-hacking. The incentives are baked in. But practical steps can move the needle. Preregistration or prespecification makes a bright line between confirmatory tests and the exploratory fishing that every good scientist does. Full reporting—effect sizes, all p-values to three decimal places, sample sizes, and a clear description of the analysis path—gives readers the context to judge robustness. When possible, blind analysis, where decisions are made without seeing which option helps a favored result, reduces unconscious nudging. And open data lets others reanalyze, catch misreports, and try alternative models. None of this requires waiting for a perfect system. It just aligns the day-to-day practice of analysis with the ideals we say we value. So here's the balanced picture to carry forward. As Head and colleagues showed, p-hacking is widespread enough to leave fingerprints you can see across entire fields. It clusters right under 0.05. It can inflate estimated effects, especially when misreported p-values sneak in. Yet across those same fields, the dominant contour of the p-curve tilts toward the tiny p-values that true effects generate. Meta-analyses, done with care, appear to pick up that signal and often blunt some of the inflation. If you're a researcher, that means two things. First, treat the exact shape of a p-curve near 0.05 as a mirror. If it bulges there, ask why, and be explicit about how you handled outliers, covariates, and stopping rules. Second, design and report so that someone else could reproduce your choices. If you're a reader or a policymaker, resist the reflex to equate a p-value less than 0.05 with importance. Look for the pattern underneath: are there many very small p-values, or does the action cluster at the line? Science will probably always have a little noise at the edges. The point isn't to pretend those edges don't exist. It's to know where they are, measure them, and keep the center of gravity on results that will hold up when the bright light of replication hits. That's what a good p-curve, and the culture that goes with it, can help us do.

If you've ever wondered why the number 0.05 shows up like a speed limit sign in so many papers, here's the story. For almost a century, scientists have leaned on a ritual called null hypothesis significance testing. You pose a null hypothesis—no effect, no relationship—run your study, and compute a p-value: the probability, if the null were true, of seeing data as extreme as yours.

Then comes the cliff. If the p-value falls below 0.05, the result gets stamped "significant" and is treated as a win. If it's above 0.05?

It's often ignored. That cutoff was meant to separate signal from noise. In practice, it shapes careers, headlines, and what even counts as knowledge.

When a single number carries that much weight, it starts to bend behavior. Head and colleagues gave a crisp name to the result: inflation bias, better known as p-hacking. You try different analysis choices—stop data collection when the graph looks good, test many outcomes and report the ones that pop, decide on outliers after a peek—and you keep the paths that land you under 0.05.

Not necessarily with malice. The incentive structure does the nudging. Journals favor "significant" results.

Funders and hiring committees count them. And there's a common confusion in the background: a small p-value can come from a tiny effect in a huge sample. It is not a measure of importance.

All of this means that the pressure to deliver a p-value less than 0.05 is both cultural and structural.

So how do you tell when a literature is signaling a real effect versus telling you what it thinks you want to hear? One surprisingly simple diagnostic is the p-curve, the distribution of reported significant p-values. Picture it this way.

If there's a true effect out there, very small p-values accumulate more than modestly small ones, so the curve leans to the right—more values of 0.01 than near 0.05. If there's no effect at all and researchers are just publishing the "wins," the p-values among those wins tend to look flatter. Now add p-hacking.

When researchers nudge near-misses over the line, you get a bump—an overabundance of p-values just below 0.05. That little cliff shapes the landscape. Crucially, the exact shape depends on two forces: whether a true effect exists and how hard people are pushing their analyses toward significance.

Head, Holman, Lanfear, Kahn, and Jennions asked a big, blunt question: across science, do we see the signal of real effects, the fingerprints of p-hacking, or both? And do the conclusions we draw from meta-analyses—our best attempt at synthesis—survive those distortions? They built a scalable toolkit to find out.

First, they text-mined p-values from open access papers across fields, separating the numbers reported in results sections from those in abstracts. Second, they reanalyzed the p-values that underlie published meta-analyses, where the questions are clearly defined and the data have already been assembled.

The core test is elegant. For evidential value, they compare how many p-values fall between zero and 0.025 to how many land between 0.025 and 0.05. More in the lower bin means a right-skewed curve and a likely real effect.

For p-hacking, they zoom in on the danger zone with a finer split: values of 0.04 to 0.045 versus values of 0.045 to 0.05. If the upper sliver is stuffed, that's a red flag. Statistically, they frame these as binomial questions: is the proportion in the upper bin different from 0.5?

They use a binomial generalized linear model to test that intercept. They also sanity-check with a sign-test variant popularized by Uri Simonsohn and colleagues. To keep the text-mined data from being dominated by prolific writers, they bootstrap one p-value per paper section a thousand times.

And when they quote proportions, the confidence intervals come from the Clopper-Pearson method, the exact one you get from binom.test in R.

Across the text-mined literature, the first message is reassuring. Pulling p-values from results sections across 14 disciplines, the proportion in the upper evidential bin—those between 0.025 and 0.05—lands at 0.257. Abstracts tell a very similar story across 10 disciplines with 0.262.

That's a right-skew, just as you'd expect if many reported effects are real. You can think of it like this: there are more very small p-values than near-threshold ones, which is hard to fake consistently without a true signal underneath.

But the second message is sobering. Zooming in on the sliver just below 0.05, the proportion tips the other way: 0.546 in results and 0.537 in abstracts. That is, there are more p-values in the 0.045 to 0.05 window than in the 0.04 to 0.045.

Exactly where you'd expect to see the fruits of garden-of-forking-paths decisions, you do. Many individual disciplines show this bump clearly. And here's the twist: when the team stripped out misreported p-values—those labeled as "p less than 0.05" that, on inspection, weren't—the p-hacking signal shrank dramatically in the discipline-wide summary.

Misreporting itself is, of course, part of the problem, but it also exaggerates how strong the p-hacking bump looks when you scan the literature at scale.

The second line of evidence focuses on the spine of our evidence base: meta-analyses. Here, Head and colleagues reassembled the p-curves from the primary data used in 12 published syntheses, each addressing a specific, pre-stated question. Again, you get that reassuring rightward lean.

The pooled upper-bin proportion is 0.202 across these datasets, and nine of the twelve individual curves show evidential value on their own. The three holdouts tended to have small sample sizes, which likely just means they lacked the statistical power to show the skew, not that the underlying effects were absent.

What about p-hacking within these meta-analytic datasets? Misreported p-values did crop up in the pile. When the team left those in, the fine-grained test near the threshold flagged a problem: the pooled proportion in the p equal to 0.045 to 0.05 window rose to 0.615 with a statistically significant test result.

Remove the misreports, and that proportion drops to 0.489 and the signal evaporates statistically. One large meta-analysis by Jiang and collaborators provides a vivid example of this contrast; the apparent p-hacking signal hinged on the misreports. The takeaway is not that p-hacking disappears inside meta-analyses.

It's that a chunk of what you detect can be an artifact of p-values that were mislabeled as "significant" in the first place.

Let's pause and translate that into plain consequences. Across fields, there is real signal—more tiny p-values than near-threshold ones—so the scientific literature isn't just smoke and mirrors. At the same time, the neighborhood right under 0.05 is crowded, which is exactly where p-hacking would park its cars.

And when you pool studies in a meta-analysis, that crowd can inflate the average effect size a bit. But because meta-analyses weight studies, the noisiest, smallest sample ones—the ones most vulnerable to p-hacking—tend to count for less. That doesn't make them immune.

It does mean synthesis can, under decent conditions, dampen the worst inflations.

Now, methods have limits, and Head and colleagues are candid about them. The two-bin evidential test is tuned to catch strong right-skews; it's less sensitive to subtle p-hacking. The hyper-focused 0.04 versus 0.045 split is great for zeroing in near the threshold, but making bins smaller boosts sensitivity at the cost of statistical power.

And text-mined p-values mash together many different questions within a discipline. A right-skew there tells you that, in aggregate, the field is finding real effects. It can't guarantee that any single, well-defined hypothesis is on solid footing.

Signals also depend on average effect sizes and study power. Underpowered domains can look less evidential even when the effects are real.

With those caveats, the toolkit they showcase—p-curves, binomial modeling of adjacent bins, bootstrap sampling to avoid overcounting prolific authors, and exact binomial intervals—does something valuable. It turns a vague worry into testable shape. If you see an overall right-skew and a bump just below 0.05, you can quantify both: how much real signal there is, and how much near-threshold inflation you're likely facing.

And because you can apply the same diagnostics to the primary data behind a meta-analysis, you can check whether synthesis itself is amplifying, damping, or surviving those pressures.

What should we do about the pressures that generate the bump in the first place? No single reform will erase p-hacking. The incentives are baked in.

But practical steps can move the needle. Preregistration or prespecification makes a bright line between confirmatory tests and the exploratory fishing that every good scientist does. Full reporting—effect sizes, all p-values to three decimal places, sample sizes, and a clear description of the analysis path—gives readers the context to judge robustness.

When possible, blind analysis, where decisions are made without seeing which option helps a favored result, reduces unconscious nudging. And open data lets others reanalyze, catch misreports, and try alternative models. None of this requires waiting for a perfect system.

It just aligns the day-to-day practice of analysis with the ideals we say we value.

So here's the balanced picture to carry forward. As Head and colleagues showed, p-hacking is widespread enough to leave fingerprints you can see across entire fields. It clusters right under 0.05.

It can inflate estimated effects, especially when misreported p-values sneak in. Yet across those same fields, the dominant contour of the p-curve tilts toward the tiny p-values that true effects generate. Meta-analyses, done with care, appear to pick up that signal and often blunt some of the inflation.

If you're a researcher, that means two things. First, treat the exact shape of a p-curve near 0.05 as a mirror. If it bulges there, ask why, and be explicit about how you handled outliers, covariates, and stopping rules.

Second, design and report so that someone else could reproduce your choices. If you're a reader or a policymaker, resist the reflex to equate a p-value less than 0.05 with importance. Look for the pattern underneath: are there many very small p-values, or does the action cluster at the line?

Science will probably always have a little noise at the edges. The point isn't to pretend those edges don't exist. It's to know where they are, measure them, and keep the center of gravity on results that will hold up when the bright light of replication hits. That's what a good p-curve, and the culture that goes with it, can help us do.

More in Decision Sciences