p-Curve and p-Hacking in Observational Research
The p-curve was supposed to be the lie detector of science. You stack up a pile of published studies, pull out their statistically significant p-values, plot the distribution, and the shape tells you the truth: are these results real, or did researchers manufacture them by fishing for significance? The tool felt like a genuine breakthrough. You didn't need to re-run the original experiments. You just needed the published numbers. And it worked — or so the field believed — because the logic was clean. Real effects produce p-values that cluster toward zero. P-hacking produces p-values that cluster just below 0.05. Different shapes, different verdicts. Then Stephan Bruns and John Ioannidis ran the simulations and found that the same curve shape which signals a true effect appeared when the true effect was exactly zero. That's the finding at the heart of their paper, published in PLOS ONE. To understand why it matters, you need to grasp both what the p-curve promised and where that promise breaks down. The p-curve concept was introduced by Simonsohn and colleagues, who proposed using the distribution of statistically significant p-values as a diagnostic for the replication crisis. The core logic goes like this: when effects are real and studies have adequate statistical power, p-values will tend to be very small, so the resulting distribution skews right — lots of values near zero, and relatively few crowed just below 0.05. P-hacking flips that picture.
If researchers are selectively reporting only the analyses that squeaked past significance, you get a spike just under 0.05 and a left-skewed distribution. The shape of the curve, in other words, carries information about the integrity of the underlying literature. Jager and Leek formalized a version of this, modeling the p-curve as a mixture of uniform and beta distributions, and estimated a 14 percent false discovery rate in the medical literature, which translates to roughly 86 percent of published findings reflecting true effects. Others used p-curve summaries to vindicate entire bodies of observational research across disciplines. The diagnostic was embraced quickly, and for understandable reasons. However, Bruns and Ioannidis identify a structural flaw specific to observational research — the kind that dominates medicine, economics, public health, and the social sciences — where the data are not generated by randomized experiments. The flaw is omitted-variable bias. When a confounder — a variable that genuinely drives the outcome and is correlated with the variable being studied — is left out of a regression model, the estimated effect gets distorted in a systematic, not random, way. Unlike experimental p-hacking, which exploits chance, omitted-variable bias produces estimates that actually converge away from the truth as sample sizes grow.
The larger your dataset, the more confidently wrong you can be. And critically, in observational research, which variables to control for is always a judgment call. The specification is always flexible. That flexibility is the attack surface. Bruns and Ioannidis model this directly with Monte Carlo simulations—computational experiments that run a data-generating process hundreds of thousands of times to see what patterns emerge. They set the true effect to exactly zero in every simulation. They then introduce omitted-variable bias by generating a confounder z that correlates with the variable of interest x at a fixed covariance of 0.2, and draw the confounder's coefficient from a uniform distribution up to some maximum. The expected maximum bias is simply 0.2 times that maximum coefficient. They run 500,000 iterations, vary sample sizes from a minimum of 50 up to maximums of 100, 1,000, 10,000, and 100,000, and test three bias strengths—those that induce maximum Pearson correlations between the outcome and predictor of 0.01, 0.05, and 0.1. Even a correlation of 0.1 is what statisticians call small, citing Cohen's conventional benchmarks. These are not dramatic, detectable biases. These are the kinds of confounding that routinely go unnoticed. The result is unambiguous. Across these simulations, even the minimal biases generate right-skewed p-curves once sample sizes are large enough. Not left-skewed.
Not spiked below 0.05. Right-skewed — the very pattern that has been interpreted as evidence of true effects. The lie detector gives the same reading for truth and lies. To drive this out of the simulation and into the real world, Bruns and Ioannidis construct an empirical demonstration using malaria prevalence and economic growth between 1960 and 1996. They start from the Sala-i-Martin dataset of 68 potential growth covariates, select 15, and use malaria prevalence in 1966 as the variable of interest. Then they do something clever: they construct a version of the growth variable from which the direct effect of malaria has been stripped out entirely. This is a null-by-construction predictor. Whatever the p-curve says about it cannot be attributed to a real effect. A typical growth regression in this literature uses seven variables — malaria plus six controls. Choosing six controls from the available fifteen yields five thousand and five possible model specifications. The dataset covers 99 countries. To avoid relying on a single sample, they draw 100 random country samples of sizes between 50 and 99 and estimate all five thousand and five models in each. That produces five hundred thousand and five hundred coefficient estimates. To visualize what this specification flexibility does, they construct a vibration plot — point estimates on one axis, transformed p-values on the other.
What you see is an effect that vibrates wildly. The malaria coefficient swings from positive to negative depending entirely on which six controls were included. Most estimates — sixty-two point six percent — are negative and statistically insignificant. However, twenty-three point four percent are negative and significant at a p-value below 0.05. That's the pool a researcher who wanted to report a malaria-growth relationship could draw from, without ever falsifying a single number. To simulate p-hacking directly, Bruns and Ioannidis allow a search across specifications and samples, stopping each time they find a negative, statistically significant estimate, until they collect one hundred thousand such results. Then, they plot the p-curve. It is right-skewed. The null variable, in the hands of a flexible analyst, produces the exact distributional signature that has been used to certify true effects in the published literature. What does this mean for the claims already on the books? Bruns and Ioannidis are explicit. The influential findings from Simonsohn and colleagues — that right-skewed p-curves signal genuine effects — are unreliable in observational settings.
Head and colleagues' inference, based on text-mined p-values across disciplines, that most published results reflect true relationships is in question. Jager and Leek's fourteen percent false discovery rate estimate for the medical literature is also in question. As Bruns and Ioannidis put it, in the presence of omitted-variable bias, "the false discovery rate could be anything, even 100 percent." That is not a rhetorical flourish; it follows directly from the simulation results. Three things can help. First, pre-registration: committing to your model specification before seeing the data removes most of the flexibility that omitted-variable p-hacking exploits. Second, empirical calibration of p-values. Schuemie and colleagues demonstrated that when you empirically calibrate significance thresholds in observational research, at least fifty-four percent of findings that passed a p-value below 0.05 become statistically insignificant. That's a majority of results that the uncalibrated threshold certified as real. Third, and most basically, stop treating p-curve analysis as a clean diagnostic in any setting where randomization is absent or compromised.
None of this is a claim that science is broken or that the replication crisis literature has been wasted effort. The p-curve is a real tool with real diagnostic power in experimental contexts where randomization holds. The contribution of Bruns and Ioannidis is narrower and more specific: the tool has a blind spot, and that blind spot covers most of the published scientific literature. In fields where you cannot randomly assign malaria, or poverty, or diet, or education — which is most fields — the shape of the p-curve is no longer a reliable verdict. It's a shape that can be produced by true effects, and it can be produced by nothing at all. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- On triangulated orbit categories
- Reporting Bias in Drug Trials Submitted to the Food and Drug Administration: Review of Publication and Presentation
- Some results on difference polynomials sharing values
- Modeling 3D Facial Shape from DNA
- A duality theorem for Willmore surfaces
- A classification theorem for nuclear purely infinite simple $C^*$-algebras