The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
In the spring of 2020, governments around the world were making enormous bets under conditions of near-total uncertainty. Should they close schools or keep them open? Shut down workplaces? Restrict internal movement? Each decision carried real costs — to children's education, to livelihoods, and to the basic texture of daily life. Governments were not making these choices one at a time, testing each before trying the next. They were stacking them, layering intervention upon intervention, often within days of each other. The central problem that Liu and colleagues set out to solve is this: when everything happens at once, how do you figure out what’s actually working? Their answer was to build the largest cross-country study of non-pharmaceutical interventions, or NPIs, attempted during the first wave of the pandemic. One hundred thirty countries and territories, thirteen categories of intervention, and six months of data, from January to June 2020. The policy data came from the Oxford COVID-19 Government Response Tracker, known as OxCGRT, which recorded government responses across one hundred seventy-eight countries. Liu and colleagues trimmed that to thirteen interventions they judged likely to affect local transmission, grouped into four broad families: internal containment and closures, international travel restrictions, economic measures like income support and debt relief, and health system actions like testing and contact tracing.
To measure transmission, the study used Rt — the time-varying reproduction number — drawn from the EpiForecasts repository. Rt is the mean number of secondary cases generated by one infected person at a given moment in time. When Rt is above one, the epidemic is growing. Below one, it’s shrinking. At the start of the pandemic, estimates for SARS-CoV-2 ranged from about two to four. The goal of every NPI on this list was to push that number down. The measurement framework sounds clean, but the reality was not. Before Liu and colleagues could run a single regression, they had to reckon with the fact that NPIs don’t arrive independently — they arrive in waves. Countries tended to escalate many measures simultaneously, particularly in March 2020, with stringency peaking around April thirteenth across most regions before declining slightly by June. To capture this, the team ran a hierarchical cluster analysis on the thirteen NPI time series using Ward’s method, tested with ten thousand bootstrap replicates. The result was stark: under their "any effort" coding — where any non-zero policy record counted as active — all thirteen NPIs collapsed into just two significant temporal clusters. They moved together. That’s the structural confounding problem in a nutshell: if school closures, workplace closures, and internal movement restrictions all happen within the same two-week window, you can’t simply look at which one correlated with a falling Rt and declare victory.
The temporal clustering also meant that lag structure mattered a great deal. There’s a delay between when a government imposes a restriction and when that restriction shows up as a change in measured Rt — because transmission events that happened before the restriction still work their way through the incubation period and into reported case counts. Liu and colleagues searched lags from one to twenty-one days and found that the best-fitting lag averaged about a week, but varied by region. East Asia and the Pacific showed lags between five and ten days. Europe showed about five days. Latin America came in below five. To handle this heterogeneity, the team ran regressions at lags of one, five, and ten days and tested two NPI coding scenarios — "any effort" and "maximum effort," the latter coding only the highest-intensity version of each policy as active. They also varied model selection criteria between AIC and BIC and ran models on both the full time series and a truncated version cutting off at peak stringency. That ensemble of specifications is the study’s core methodological answer to the attribution problem: rather than trusting any single model, they looked for findings that held across the range. So what did hold? The strongest finding is simple. Two NPIs showed consistent, robust associations with reduced Rt across all model specifications: complete school closure and restrictions on internal movement.
Liu and colleagues describe this result as unequivocal. It survived different lag assumptions, different coding scenarios, and different selection criteria. If you want to know what the data most clearly supported, it’s those two. Everything else is more conditional, and understanding those conditions is where the real policy texture lies. Three interventions — workplace closure, income support, and debt or contract relief — showed strong evidence of association with reduced Rt, but only under the any-effort scenario. In other words, the signal appeared when these policies were first initiated or present at any level, but there was no evidence that pushing them to maximum intensity produced larger additional reductions. Two other NPIs — cancellation of public events and restrictions on gatherings — showed strong evidence only at maximum intensity. Restrictions on gatherings of one thousand or more people? Not effective in these models. Restrictions on gatherings of fewer than ten? That’s where the signal appeared. The threshold mattered enormously. Then there’s the inconclusive tier, and this is where the clustering problem bites hardest. Stay-at-home requirements, public information campaigns, public transport closure, international travel controls, testing policies, and contact tracing all produced weak or inconsistent evidence of association with Rt reductions. Liu and colleagues are explicit that this does not mean these interventions are ineffective.
It means the statistical signal couldn’t be cleanly separated from the temporally clustered policies that deployed alongside them. When a cluster of interventions coincides with a falling Rt, the data can’t always tell you which members of that cluster deserve the credit. Contact tracing, for instance, was often rolled out alongside school closures and movement restrictions — and in that context, isolating its independent contribution becomes analytically very difficult. Two numbers frame the scope of this ambiguity: one hundred thirty countries analyzed and thirteen interventions tested. The breadth is real, but so is the fundamental limitation of observational ecological data. The study’s model estimates Rt for a country on a given day as a country-specific baseline level, plus the combined contribution of all active NPI indicators on that day, plus error. That structure can reveal associations, but it cannot establish clean causation. The model doesn’t propagate uncertainty in the Rt estimates themselves, doesn’t capture interactions between NPIs, and relies on assumptions about roughly constant case ascertainment across the March to June 2020 window — assumptions that are at best approximate in a period of rapidly expanding testing.
The OxCGRT database itself carries limitations that the authors acknowledge. Many specific measures — face-covering mandates, for instance — aren’t independently coded. A policy coded as "comprehensive" contact tracing in one country might look very different in operational terms from the same code in another. The ordinal intensity scales that underlie the any-effort and maximum-effort coding scenarios don’t guarantee that step two is meaningfully different from step one. None of this undermines the study’s genuine contribution. A panel of one hundred thirty countries using consistent data across six months, tested against multiple model specifications, with explicit attention to temporal clustering — that’s the broadest empirical evaluation of NPI effectiveness produced in the first wave. The finding that school closures and internal movement restrictions consistently reduced Rt is meaningful precisely because it survived so many different analytical choices. Additionally, the finding that gathering restrictions only worked at maximum intensity carries a direct policy implication: a rule banning events of one thousand or more people was not, in these data, detectably bending the curve. Banning gatherings of ten or more was. What this study cannot answer — and what Liu and colleagues are honest about — are the questions that follow. Which NPIs should a government reach for first? Which should it be slowest to relax?
How do effects change as a population’s behavior adapts over months? Those questions require complementary evidence, different study designs, and longer time horizons. This dataset covers only the first six months of the pandemic, before widespread immunity, before vaccines, and before the behavioral and institutional learning that came later. What it does capture is the first global experiment in epidemic control by social policy — thousands of decisions made under pressure, affecting hundreds of millions of people. Understanding which of those decisions actually drove Rt below one isn’t just an academic exercise. It’s the foundation for making better choices the next time. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Estimating the infection and case fatality ratio for coronavirus disease (COVID-19) using age-adjusted data from the outbreak on the Diamond Princess cruise ship, February 2020
- Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
- Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School
- Early dynamics of transmission and control of COVID-19: a mathematical modelling study