Do Pressures to Publish Increase Scientists' Bias? An Empirical Support from US States Data

Daniele FanelliView original
OverviewBalancedharper voice
You have run the experiment. Six months of work, maybe a year. The results came back, and they didn't support your hypothesis. Now you are sitting with a null finding — technically valid, scientifically informative, and almost impossible to publish. The grant renewal is in eight months. The tenure clock is running. What do you do? That moment of pressure is exactly what Daniele Fanelli decided to measure. The starting premise of Fanelli's study is that academic careers increasingly depend on publication volume and the prestige of the journals where those publications appear. In that environment, a well-documented fact becomes a structural problem: papers reporting positive results — findings that support the hypothesis under test — are more likely to be published, more likely to be cited, and more likely to land in high-impact journals. Negative results, those that fail to support a hypothesis, tend to stay in file drawers. The file-drawer effect is one of two main mechanisms that translate career pressure into a skewed published record. The second is HARKing — Hypothesizing After the Results are Known — where researchers reframe their hypotheses after the fact so that whatever came out of the experiment appears to confirm something. Add selective reporting of measures and post-hoc tweaking of analyses, and you have a system that, even without any outright fraud, could quietly fill the literature with results that skew positive. The problem is well known in qualitative terms. But Fanelli wanted to measure it empirically. And to do that, he needed a clever proxy for competitive pressure. His solution was geography. If publish-or-perish pressure is real and varies in intensity across academic environments, then US states offer a natural experiment. States with higher per-capita academic output — more papers per science and engineering doctorate holder, measured by National Science Foundation data — represent more competitive, more productive institutional environments. The question becomes: do papers coming out of those high-pressure states report positive results at a higher rate? To test this, Fanelli drew a large random sample from the Essential Science Indicators database, searching for papers containing the phrase "test the hypothesis" or close variants, published between two thousand and two thousand seven, with a corresponding author in the United States. The final sample was one thousand three hundred sixteen papers, representing every state except Delaware. Each paper's abstract or full text was read to classify the outcome as positive — full or partial support for the hypothesis — or negative. Crucially, this coding was done blind to the author's address, and a separate rater, given only brief instructions, agreed with the primary coding in 18 out of 20 test cases, confirming that the classification was replicable. The statistical workhorse was logistic regression — a model that expresses the log-odds of a paper reporting a positive result as a function of predictors. The main predictor of interest was the state's per-capita publication rate. The models also controlled for state per-capita research and development expenditure, study methodology, whether multiple hypotheses were tested, and discipline. The range of outcomes across states alone is striking. The proportion of positive results by state ran from twenty-five percent at the low end to one hundred percent at the high end, with a mean of eighty-two percent. That spread is not random noise — and the regression confirms it isn't. In the baseline model, per-capita academic productivity significantly predicted whether a paper reported a positive result: a slope of one point thirty-eight, with an odds ratio of just under four. But here is the finding that sharpens the story considerably. When Fanelli added a control for state per-capita research and development expenditure, the productivity effect got larger, not smaller. The odds ratio jumped to over fourteen. And research and development spending itself showed a negative, non-significant trend — meaning more money didn't produce more positive results; more competitive pressure did. Let that sit for a moment. You might expect that richer states, with better-funded labs and more resources, would simply run better experiments and legitimately confirm more hypotheses. The data don't support that story. Resources going up did not push positive rates up. Competitive publication pressure going up did. When Fanelli added controls for methodological characteristics, the slope for per-capita productivity came in at two point fifty-nine, with an odds ratio of twelve point twenty-nine and a ninety-five percent confidence interval of two point zero two to eighty-seven point thirty-seven. When he swapped in discipline controls instead, the slope was two point fifty-one, odds ratio twelve point twenty-nine. Both specifications told the same story: the association between a state's publication intensity and the probability of a positive result is real, consistent across model specifications, and actually grows when you account for factors that might otherwise explain it away. Fanelli also tested whether the effect was stronger in some disciplines than others by adding an interaction term. Overall, that interaction didn't improve model fit — suggesting the pattern cuts across fields rather than being driven by one or two. Two disciplines did show notably stronger associations in the data: Pharmacology and Toxicology, and Neuroscience and Behavior. Fanelli flags these as possibly due to chance, given the multiple comparisons involved. Now, Fanelli is careful here, and the carefulness matters. The most flattering alternative explanation is that researchers in high-productivity states are working in more prestigious institutions with better equipment, sharper colleagues, and stronger experimental designs — and so they genuinely get more true positive results. He calls this an unavoidable confounding factor. The study cannot fully rule it out. But he lays out why the pattern is harder to explain as pure quality. The productivity effect growing stronger after controlling for research and development expenditure is the key signal. If superior resources explained the results, controlling for spending should absorb the effect. It doesn't. If anything, spending and positive rates move in opposite directions. And the breadth of the pattern across disciplines — confirmed by the non-significant interaction term — argues against the idea that a handful of elite universities in a particular field are driving the whole thing. There is also the publication-bias alternative: editors and reviewers may favor authors from prestigious institutions, which happen to cluster in high-productivity states. Fanelli acknowledges this but notes that such editorial favoritism could plausibly work in either direction when it comes to outcome favorability, and the available evidence doesn't cleanly support it as the primary driver. The honest conclusion Fanelli reaches is this: the data support the hypothesis that competitive academic environments increase not just scientists' productivity but their bias. That is evidence, not proof. Institutional prestige cannot be fully excluded. But the pattern is real, consistent, and directionally clear. What does this mean at scale? If competitive pressure systematically inflates the rate of positive results across an entire research literature, that literature becomes unreliable in ways that are hard to detect from inside it. A single reader cannot tell, from a paper's methods section, whether a null finding was tucked away and the analysis was reframed to produce something publishable. But in aggregate, across thousands of papers, the signal shows up in the data Fanelli assembled. Meta-analyses — systematic reviews that pool results across dozens or hundreds of studies — inherit this bias. Policy built on those meta-analyses inherits it too. Fanelli notes that the same dynamic almost certainly operates in other countries where academic competition and publication pressure are high. The US states here are a case study in a global phenomenon. The incentive structure is not uniquely American. The core irony is worth sitting with. The system of peer-reviewed publication was designed to filter knowledge — to expose findings to scrutiny and let the best evidence accumulate. But when the reward for publishing is high enough, and the penalty for null results is steep enough, that system begins producing something subtly different from knowledge. Not lies, exactly. Not mostly fraud. But a literature tilted, measurably and consistently, toward results that confirm rather than disconfirm. A record that looks more certain than it is. And the more competitive the environment, the more pronounced that tilt becomes — which means the most productive corners of science, the ones we look to first, may be the most systematically optimistic. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

You have run the experiment. Six months of work, maybe a year. The results came back, and they didn't support your hypothesis. Now you are sitting with a null finding — technically valid, scientifically informative, and almost impossible to publish. The grant renewal is in eight months. The tenure clock is running. What do you do? That moment of pressure is exactly what Daniele Fanelli decided to measure. The starting premise of Fanelli's study is that academic careers increasingly depend on publication volume and the prestige of the journals where those publications appear. In that environment, a well-documented fact becomes a structural problem: papers reporting positive results — findings that support the hypothesis under test — are more likely to be published, more likely to be cited, and more likely to land in high-impact journals. Negative results, those that fail to support a hypothesis, tend to stay in file drawers. The file-drawer effect is one of two main mechanisms that translate career pressure into a skewed published record. The second is HARKing — Hypothesizing After the Results are Known — where researchers reframe their hypotheses after the fact so that whatever came out of the experiment appears to confirm something. Add selective reporting of measures and post-hoc tweaking of analyses, and you have a system that, even without any outright fraud, could quietly fill the literature with results that skew positive.

The problem is well known in qualitative terms. But Fanelli wanted to measure it empirically. And to do that, he needed a clever proxy for competitive pressure. His solution was geography. If publish-or-perish pressure is real and varies in intensity across academic environments, then US states offer a natural experiment. States with higher per-capita academic output — more papers per science and engineering doctorate holder, measured by National Science Foundation data — represent more competitive, more productive institutional environments. The question becomes: do papers coming out of those high-pressure states report positive results at a higher rate? To test this, Fanelli drew a large random sample from the Essential Science Indicators database, searching for papers containing the phrase "test the hypothesis" or close variants, published between two thousand and two thousand seven, with a corresponding author in the United States. The final sample was one thousand three hundred sixteen papers, representing every state except Delaware. Each paper's abstract or full text was read to classify the outcome as positive — full or partial support for the hypothesis — or negative. Crucially, this coding was done blind to the author's address, and a separate rater, given only brief instructions, agreed with the primary coding in 18 out of 20 test cases, confirming that the classification was replicable.

The statistical workhorse was logistic regression — a model that expresses the log-odds of a paper reporting a positive result as a function of predictors. The main predictor of interest was the state's per-capita publication rate. The models also controlled for state per-capita research and development expenditure, study methodology, whether multiple hypotheses were tested, and discipline. The range of outcomes across states alone is striking. The proportion of positive results by state ran from twenty-five percent at the low end to one hundred percent at the high end, with a mean of eighty-two percent. That spread is not random noise — and the regression confirms it isn't. In the baseline model, per-capita academic productivity significantly predicted whether a paper reported a positive result: a slope of one point thirty-eight, with an odds ratio of just under four. But here is the finding that sharpens the story considerably. When Fanelli added a control for state per-capita research and development expenditure, the productivity effect got larger, not smaller. The odds ratio jumped to over fourteen. And research and development spending itself showed a negative, non-significant trend — meaning more money didn't produce more positive results; more competitive pressure did.

Let that sit for a moment. You might expect that richer states, with better-funded labs and more resources, would simply run better experiments and legitimately confirm more hypotheses. The data don't support that story. Resources going up did not push positive rates up. Competitive publication pressure going up did. When Fanelli added controls for methodological characteristics, the slope for per-capita productivity came in at two point fifty-nine, with an odds ratio of twelve point twenty-nine and a ninety-five percent confidence interval of two point zero two to eighty-seven point thirty-seven. When he swapped in discipline controls instead, the slope was two point fifty-one, odds ratio twelve point twenty-nine. Both specifications told the same story: the association between a state's publication intensity and the probability of a positive result is real, consistent across model specifications, and actually grows when you account for factors that might otherwise explain it away. Fanelli also tested whether the effect was stronger in some disciplines than others by adding an interaction term. Overall, that interaction didn't improve model fit — suggesting the pattern cuts across fields rather than being driven by one or two. Two disciplines did show notably stronger associations in the data: Pharmacology and Toxicology, and Neuroscience and Behavior. Fanelli flags these as possibly due to chance, given the multiple comparisons involved.

Now, Fanelli is careful here, and the carefulness matters. The most flattering alternative explanation is that researchers in high-productivity states are working in more prestigious institutions with better equipment, sharper colleagues, and stronger experimental designs — and so they genuinely get more true positive results. He calls this an unavoidable confounding factor. The study cannot fully rule it out. But he lays out why the pattern is harder to explain as pure quality. The productivity effect growing stronger after controlling for research and development expenditure is the key signal. If superior resources explained the results, controlling for spending should absorb the effect. It doesn't. If anything, spending and positive rates move in opposite directions. And the breadth of the pattern across disciplines — confirmed by the non-significant interaction term — argues against the idea that a handful of elite universities in a particular field are driving the whole thing. There is also the publication-bias alternative: editors and reviewers may favor authors from prestigious institutions, which happen to cluster in high-productivity states. Fanelli acknowledges this but notes that such editorial favoritism could plausibly work in either direction when it comes to outcome favorability, and the available evidence doesn't cleanly support it as the primary driver.

The honest conclusion Fanelli reaches is this: the data support the hypothesis that competitive academic environments increase not just scientists' productivity but their bias. That is evidence, not proof. Institutional prestige cannot be fully excluded. But the pattern is real, consistent, and directionally clear. What does this mean at scale? If competitive pressure systematically inflates the rate of positive results across an entire research literature, that literature becomes unreliable in ways that are hard to detect from inside it. A single reader cannot tell, from a paper's methods section, whether a null finding was tucked away and the analysis was reframed to produce something publishable. But in aggregate, across thousands of papers, the signal shows up in the data Fanelli assembled. Meta-analyses — systematic reviews that pool results across dozens or hundreds of studies — inherit this bias. Policy built on those meta-analyses inherits it too. Fanelli notes that the same dynamic almost certainly operates in other countries where academic competition and publication pressure are high. The US states here are a case study in a global phenomenon. The incentive structure is not uniquely American.

The core irony is worth sitting with. The system of peer-reviewed publication was designed to filter knowledge — to expose findings to scrutiny and let the best evidence accumulate. But when the reward for publishing is high enough, and the penalty for null results is steep enough, that system begins producing something subtly different from knowledge. Not lies, exactly. Not mostly fraud. But a literature tilted, measurably and consistently, toward results that confirm rather than disconfirm. A record that looks more certain than it is. And the more competitive the environment, the more pronounced that tilt becomes — which means the most productive corners of science, the ones we look to first, may be the most systematically optimistic. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Decision Sciences