Replication, Communication, and the Population Dynamics of Scientific Discovery
Most published scientific findings are false. Not just most bad science — most published science, full stop. That's not a cynical opinion. It's the conclusion that John Ioannidis reached in two thousand five, and it's the provocation that Richard McElreath and Paul Smaldino decided to take seriously using mathematics. Here’s what their model reveals: one of the field's favorite proposed fixes — publishing everything, including the failed replications — can actually make things worse. The problem, McElreath and Smaldino argue, isn't just that individual studies fail. It's that the field has been debating remedies without a shared, formal framework. Everyone agrees that replication matters, that publication bias distorts the record, and that statistical power is too often too low — but without a model that holds all of these variables simultaneously, the debates remain verbal and imprecise. So they built one. The framework is a population dynamics model — an approach borrowed from ecology and applied to scientific belief. The core idea is simple but powerful: instead of treating each hypothesis as an isolated claim, treat hypotheses as populations that accumulate evidence over time. Every hypothesis carries a tally, which is the difference between its published positive findings and its published negative ones.
A positive result pushes the tally up; a negative one pulls it down. Hypotheses drift through tally space as studies are done and results are communicated. The question the model is built to answer is: under what conditions does this system actually converge on truth? The parameters that govern the system are few but consequential. The base rate, b, is the probability that a novel hypothesis is actually true before anyone tests it. McElreath and Smaldino stress that intuitions about b tend to be wildly optimistic. The range they illustrate spans from b below one in ten thousand for genome-wide association searches to b around one-half for something like predicting a presidential election. The false positive rate, alpha, is the probability a false hypothesis produces a positive finding anyway. And power, one minus beta, is the probability that a true hypothesis produces a positive finding. The model assumes power exceeds the false positive rate, which is a minimal condition for science to work at all. Communication adds another layer. Not every finding gets published. The model tracks three probabilities: whether a novel negative finding gets communicated, whether a negative replication gets communicated, and whether a positive replication gets communicated. Changing these probabilities alters how hypotheses flow through the tally system — and that's where the surprising results live.
The first big result concerns replication. In the model, replication acts like a ratchet. True hypotheses tend to accumulate positive tallies over time; false ones tend to drift downward. So higher tallies become enriched for truth. Under optimistic parameters — a base rate of 0.1, power of 0.8, and a false positive rate of 0.05 — a single positive replication, bringing a hypothesis to a tally of two, is often enough to make it more likely true than false. That's genuinely encouraging. But in a pessimistic regime — a base rate of one in a thousand, power of 0.6, and a false positive rate of 0.1 — a tally of five or more may be required before the majority of hypotheses at that level are actually true. The ratchet works, but the teeth wear down exactly when you need them most. Low base rates, low power, and high false positive rates all conspire to make replication less efficient. You can replicate all you want, but if those same conditions apply to the replications, you’re turning a worn ratchet. Here’s something the model reveals about replication that cuts against conventional wisdom: the replication rate — how often researchers choose to replicate versus pursue novel hypotheses — has surprisingly little impact on precision at a given tally. Replication mainly changes how quickly hypotheses reach high tallies, not how trustworthy those tallies are when they get there. That's a subtle but important distinction.
Now for the counterintuitive finding about publication bias. The standard story is that publication bias is the villain — positive results get published, null results disappear into the file drawer, and the literature skews toward false positives. McElreath and Smaldino show this is too simple. Their analogy is what they call epistemological chromatography. Think of the tally system as a separation column: true and false hypotheses enter mixed at the bottom, and replication acts as the solvent that pulls them apart. The resolution of that column — how cleanly it separates truth from falsehood — depends critically on which findings get communicated. Here’s the key distinction the model draws. When the base rate is very low, suppressing novel negative findings — the initial nulls that never make it to publication — can actually improve precision. The reason is that when b is tiny, most initial studies of new ideas are chasing false positives. Publishing those null results would flood the literature with confirmed negatives on mostly false hypotheses, creating noise. Screening them out keeps the literature tilted toward the small fraction of initial positives that are more likely to reflect real effects. So in low base rate fields, some version of the file drawer problem may be doing inadvertent epistemic work.
But — and this is critical — the same logic does not apply to negative replications. Suppressing negative replication results consistently reduces precision. Communicating failed replications aids discovery even when those replications have lower statistical power than the original studies. The two communication channels behave completely differently: screening initial nulls can help, but suppressing replication failures always hurts. Publication bias isn't a single phenomenon. It's at least two distinct phenomena, and conflating them leads to the wrong prescriptions. A concrete example from the model makes this tangible. Using a base rate of one in a thousand, replication rate of 0.2, power of 0.8, and a false positive rate of 0.05 — when both positive and negative replications are communicated, hypotheses with tallies of three or more are true more than 80 percent of the time, and more than half of all published true hypotheses have reached a tally of at least three. The separation column is working. Suppressing positive replications — publishing only 20 percent of them — raises precision at high tallies but reduces sensitivity, meaning fewer true hypotheses reach the top of the column. Suppressing negative replications degrades the whole system.
McElreath and Smaldino extend the model in two directions that bring it closer to how science actually works. First, targeted replication: what happens when researchers preferentially replicate hypotheses that already have positive tallies, rather than picking randomly? Targeting boosts sensitivity — more true hypotheses make it to high precision tallies — while leaving precision at those tallies largely unchanged. The logic is intuitive once you see it: true hypotheses diffuse upward faster than false ones, so focusing replication on intermediate tallies means you're disproportionately retesting real effects and pushing them higher. Targeting concentrates scarce resources where they convert best. Second, the model allows different research groups to operate with different levels of quality — different power and different false positive rates. This is the realistic scenario. What they find is that elevated false positive rates in initial studies cascade through the entire population of hypotheses, making the overall system harder to clean up through replication. High false positives at the entry point poison the tally distribution in a way that even high-quality replication struggles to repair. This connects directly to targeted replication: if you can't control entry-level quality across the field, targeting your replications toward the most credible signals is a partial substitute.
What does the model actually prescribe? Two levers dominate everything else. Reduce the false positive rate — through pre-registration, larger samples, and stricter significance thresholds. And raise the base rate of true hypotheses under test — through better theory, stronger prior knowledge, and fewer speculative long shots. The model shows that even modest improvements in both dimensions compound into substantially better precision across the literature. Replication helps, but it helps more when these upstream conditions are better. The deeper contribution is a shared, quantitative language for debates that have too often stayed at the level of rhetoric. When someone argues that we need more replication, the model asks: what base rate are you assuming? What’s your false positive rate? The answer changes the prescription. Science isn't broken — but it behaves like a population system, with its own dynamics, its own drift, and its own failure modes. McElreath and Smaldino's model doesn't fix those failure modes. It makes them legible. And that's where any fix has to start. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Landscapes with Araucaria in South America: evidence for a cultural dimension
- Historical Human Footprint on Modern Tree Species Composition in the Purus-Madeira Interfluve, Central Amazonia
- The relative importance of religion and education on university students’ views of evolution in the Deep South and state science standards across the United States
- Lethal Interpersonal Violence in the Middle Pleistocene
- Art therapy is associated with sustained improvement in cognitive function in the elderly with mild neurocognitive disorder: findings from a pilot randomized controlled trial for art therapy and music reminiscence activity versus usual care
- How People Domesticated Amazonian Forests