High Contagiousness and Rapid Spread of Severe Acute Respiratory Syndrome Coronavirus 2

Steven Sanche, Yen Ting Lin, Chonggang Xu, Ethan Romero-Severson, Nicolas Hengartner, Ruian KeView original
OverviewBalancedalloy voice
Picture the first weeks of 2020. News trickles out of Wuhan. Hospitals fill. Flights leave. And the question everyone in public health is asking is both simple and brutal: how fast is this thing growing? Not in a hand-wavy way, but in a way you can plug into a forecast, staff a hospital, or time a lockdown. Sanche and colleagues set out to answer that by doing something clever. Instead of staring at messy, rapidly changing case counts inside Wuhan, they watched the virus's footprints outside it. They tracked when and where the first infected travelers showed up and tied those footprints back to how big the outbreak inside Wuhan had to be to create that pattern. The backbone of their approach is easy to say out loud, and that's the point. Assume the total number of infected people in Wuhan—symptomatic or not—was growing exponentially. If you start a clock at a theoretical time when there's just one infection and then keep multiplying by a fixed factor each day, you get a curve defined by two numbers: the growth rate and that starting time. They call the total infected I-star, the growth rate r, and the start time t-zero. Then they link that internal growth to the odds an infected person gets on a bus or a train and lands in another province. If you know when the very first known infected person arrived in, say, Guangdong, and you know how many people were traveling there from Wuhan each day, you can work backwards to how big Wuhan's infected pool had to be on the day that traveler left. To make it concrete, they pulled together two unusual data streams. First, one hundred thirty-seven detailed case reports from across China—and three from outside China—each with a departure date from Wuhan. Second, anonymized mobility data from the Baidu Migration server that tracks daily traveler flows out of Wuhan to every province before the city locked down on January 23. The headline from that mobility feed is stark: between roughly forty thousand and one hundred forty thousand people left Wuhan for other provinces each day in the run-up to the lockdown. That's a lot of seeds for a lot of places. Using the earliest confirmed arrival in each of twenty-six provinces, the "first-arrival" model asks: what growth rate inside Wuhan makes those arrival times most likely? The answer lands high. They infer an r of about 0.29 per day, with a plausible range from 0.21 to 0.37. Translate that into something your gut understands and it's a doubling time of about 2.4 days, with uncertainty stretching from just under two days to just over three. The model also puts the start of sustained exponential growth around December 20, 2019, give or take a week. Once you have r and t-zero, you can read off how many infections there likely were in Wuhan on specific dates. Sanche and colleagues estimate that around January 18, the infected pool in Wuhan was already in the low thousands—about four thousand one hundred—and by January 23, the day of the lockdown, it had climbed to roughly eighteen thousand seven hundred. Left unchecked, that trajectory would have produced on the order of a few hundred thousand infections by the end of January. Those are not case counts. Those are estimates of total infections, including people who hadn't yet shown up at a hospital. But they didn't stop at one model. Because exported cases didn't just arrive first in each province; they kept coming. The second approach fixes Wuhan's internal growth as that same exponential engine—again, I-star equals a constant times e to the r times t—and then simulates what happens to infected travelers once they leave using a stochastic Susceptible-Exposed-Infectious-Removed process. SEIR just means people are Susceptible, then Exposed but not yet infectious, then Infectious, and then Removed—either recovered or isolated. The "stochastic" part acknowledges real life: who travels when, who gets hospitalized when, who gets confirmed when—all have randomness baked in. This hybrid case-count model is fitted to daily new cases reported in provinces outside Hubei during a tight window, from January 19 to January 26, when most of those cases were tied to Wuhan exposure. They use the same Baidu travel volumes and they wire in timing distributions recovered from the case reports for steps like symptom onset to hospitalization to confirmation. It's a different lens on the same question. The payoff is that the two lenses converge. The case-count model gives a best-fit growth rate of 0.30 per day, with a ninety-five percent range from 0.26 to 0.34, and pegs the theoretical start of exponential growth a few days earlier, around December 16, 2019. That's the same story told two ways: rapid doubling, mid-December takeoff. And if you're worried those numbers might just be artifacts of detection biases, they ran a simple cross-check: deaths, which are generally harder to miss, were themselves growing exponentially in late January at around 0.22 to 0.27 per day. Slower than cases, as you'd expect given the lag to death, but marching to the same drumbeat. There's one more layer here that matters for policy: sensitivity. They pushed on twenty-three different "what if" scenarios to see how assumptions about surveillance change the answer. Suppose a constant fraction of infections was never detected outside Hubei. That shifts t-zero earlier—you had more infections than you thought—but it barely nudges the growth rate. Suppose surveillance outside Hubei ramped up over January. That can pull the estimated r down. In their most conservative setup, the growth rate dips to about 0.21 per day. High either way. The point is, under very different detection stories, you still need a rapidly growing outbreak inside Wuhan to explain what the rest of China was seeing. Now, growth rate is one thing. Transmissibility is another. The bridge between them is the serial interval—the time between when one person gets infected and when the people they infect get infected—which depends on how long people incubate the virus and how long they're infectious. Using early Wuhan estimates that put the serial interval around seven to eight days, Sanche and colleagues back out a basic reproductive number, R-naught, of about 5.8. The ninety-five percent interval runs from 4.4 to 7.7. Widen that serial-interval range to six to nine days and the median barely moves—to about 5.7—with an uncertainty window from 3.8 to 8.9. That's uncomfortably high. Much higher than the early, low-twos R-naught values that floated through the first headlines. Those R-naught values come with a quiet corollary. If each infected person infects, on average, between four and eight others in a fully susceptible population, single interventions won't cut it. You need layers: surveillance to catch cases earlier, quarantine and contact tracing to box in chains of transmission, and strong social distancing to pull the average number of contacts way down. That's not spin; it's arithmetic. In a classic herd-immunity calculation, the fraction of the population that needs to be immune or insulated from infection is one minus one over R-naught. When that denominator is five or six, the target for control moves far out. Along the way, they also extract the time constants of the disease from outside Hubei case reports. The average incubation period—the time from exposure to symptoms—lands around 4.2 days, with a plausible range from 3.5 to 5.1. There's a telling shift in behavior as awareness grew: the time from symptom onset to hospitalization dropped from about 5.5 days before January 18 to 1.5 days after, a change that's statistically robust. The time from symptom onset to death averages about 16.1 days. Those numbers set the pace of the clinical course and explain why case counts and death counts look like offset waves. Hospital stays track with that story. On average, people discharged recovered after about eleven and a half days in the hospital, while those who died did so after roughly eleven days of hospitalization. That's the human timeline behind the transmission math, and it frames why health systems felt slammed: fast-growing infections layered on multi-week clinical arcs. One subtle but important assumption underlies both models: exported cases were detected perfectly. That's a strong claim in a chaotic moment. Sensitivity checks help. If, in reality, a fixed chunk of infected travelers was missed, the models respond by pushing the onset of exponential growth earlier, but the slope of that growth—the daily r—stays about the same. If detection was improving over time, some of what looks like exponential growth in exported case counts is just better surveillance. That's why their worst-case growth rate is lower, at about 0.21 per day. Even then, the doubling time is only about three days. Not exactly reassuring. It's worth pausing on how these methods read the epidemic through motion. The Baidu data are not lab tests or case forms. They're anonymized traces of where people went. Tying those flows to when the very first infected travelers surfaced in each province turns mobility into a measurement tool. In parallel, treating Wuhan's internal growth deterministically while letting the exported infections play out stochastically matches the reality that big populations obey smooth curves, while the first few cases in a new place are jumpy and discrete. As a cross-check, the two models don't just agree with themselves; they align with death trajectories and with independent lines of evidence on generation times. One more uncomfortable detail emerges in their broader checks: perhaps a fifth of transmission could have been coming from unidentified cases. If that's true, contact tracing alone can't chase it all down. It pushes you toward population-level tools—closing gathering places, widespread masking, workplace changes—that don't care whether someone has been named a case yet. And it tightens the clock. When your doubling time is two to three days, every day you wait multiplies the problem. Why does any of this history lesson matter now? Because it shows how to read an outbreak even when your central data are smoky. Watch the periphery. Measure movement. Use simple equations you can explain at a whiteboard. And stress-test your conclusions against the biases you can't avoid. Sanche and colleagues show that under a wide range of assumptions, the early Wuhan outbreak was on a trajectory that would swamp hospitals unless it was met with early, layered controls. That conclusion isn't just about one city in one month. It's about the physics of fast epidemics. When you turn the math back into policy, the message is blunt. With growth rates around 0.21 to 0.30 per day and R-naught hovering near six under plausible intervals, you need surveillance that's faster than people's social lives, quarantine and tracing that run on tight clocks, and social distancing strong enough to slash contact rates quickly. Wuhan's intense pre-lockdown mobility—those forty thousand to one hundred forty thousand daily departures—explains how sparks landed across China. The models explain why dampening those sparks needed to happen everywhere at once. There are caveats, and they're honest ones. Early case reports tilted toward severe cases. Detection outside Hubei was assumed to be perfect, which it wasn't. Parameter choices—how long people incubate, how long they're infectious—nudge R-naught up or down. But the central picture holds across the tests they ran: rapid exponential growth, early December takeoff, high transmissibility. If you want a coda, it's this. The models were simple by design. I-star equals a constant times an exponential. A clean SEIR engine for exported cases. Baidu's travel counts as a backbone. In a blizzard of uncertainty, simplicity was a feature, not a bug, because it let the team tie timing, movement, and growth into a story you can both believe and act on. And it set a standard for how to infer what matters most in the early days of a new pathogen: how fast it's growing, how widely it can spread, and how quickly you need to move to beat it.

Picture the first weeks of 2020. News trickles out of Wuhan. Hospitals fill.

Flights leave. And the question everyone in public health is asking is both simple and brutal: how fast is this thing growing? Not in a hand-wavy way, but in a way you can plug into a forecast, staff a hospital, or time a lockdown.

Sanche and colleagues set out to answer that by doing something clever. Instead of staring at messy, rapidly changing case counts inside Wuhan, they watched the virus's footprints outside it. They tracked when and where the first infected travelers showed up and tied those footprints back to how big the outbreak inside Wuhan had to be to create that pattern.

The backbone of their approach is easy to say out loud, and that's the point. Assume the total number of infected people in Wuhan—symptomatic or not—was growing exponentially. If you start a clock at a theoretical time when there's just one infection and then keep multiplying by a fixed factor each day, you get a curve defined by two numbers: the growth rate and that starting time.

They call the total infected I-star, the growth rate r, and the start time t-zero. Then they link that internal growth to the odds an infected person gets on a bus or a train and lands in another province. If you know when the very first known infected person arrived in, say, Guangdong, and you know how many people were traveling there from Wuhan each day, you can work backwards to how big Wuhan's infected pool had to be on the day that traveler left.

To make it concrete, they pulled together two unusual data streams. First, one hundred thirty-seven detailed case reports from across China—and three from outside China—each with a departure date from Wuhan. Second, anonymized mobility data from the Baidu Migration server that tracks daily traveler flows out of Wuhan to every province before the city locked down on January 23.

The headline from that mobility feed is stark: between roughly forty thousand and one hundred forty thousand people left Wuhan for other provinces each day in the run-up to the lockdown. That's a lot of seeds for a lot of places.

Using the earliest confirmed arrival in each of twenty-six provinces, the "first-arrival" model asks: what growth rate inside Wuhan makes those arrival times most likely? The answer lands high. They infer an r of about 0.29 per day, with a plausible range from 0.21 to 0.37.

Translate that into something your gut understands and it's a doubling time of about 2.4 days, with uncertainty stretching from just under two days to just over three. The model also puts the start of sustained exponential growth around December 20, 2019, give or take a week.

Once you have r and t-zero, you can read off how many infections there likely were in Wuhan on specific dates. Sanche and colleagues estimate that around January 18, the infected pool in Wuhan was already in the low thousands—about four thousand one hundred—and by January 23, the day of the lockdown, it had climbed to roughly eighteen thousand seven hundred. Left unchecked, that trajectory would have produced on the order of a few hundred thousand infections by the end of January.

Those are not case counts. Those are estimates of total infections, including people who hadn't yet shown up at a hospital.

But they didn't stop at one model. Because exported cases didn't just arrive first in each province; they kept coming. The second approach fixes Wuhan's internal growth as that same exponential engine—again, I-star equals a constant times e to the r times t—and then simulates what happens to infected travelers once they leave using a stochastic Susceptible-Exposed-Infectious-Removed process.

SEIR just means people are Susceptible, then Exposed but not yet infectious, then Infectious, and then Removed—either recovered or isolated. The "stochastic" part acknowledges real life: who travels when, who gets hospitalized when, who gets confirmed when—all have randomness baked in.

This hybrid case-count model is fitted to daily new cases reported in provinces outside Hubei during a tight window, from January 19 to January 26, when most of those cases were tied to Wuhan exposure. They use the same Baidu travel volumes and they wire in timing distributions recovered from the case reports for steps like symptom onset to hospitalization to confirmation. It's a different lens on the same question.

The payoff is that the two lenses converge. The case-count model gives a best-fit growth rate of 0.30 per day, with a ninety-five percent range from 0.26 to 0.34, and pegs the theoretical start of exponential growth a few days earlier, around December 16, 2019. That's the same story told two ways: rapid doubling, mid-December takeoff.

And if you're worried those numbers might just be artifacts of detection biases, they ran a simple cross-check: deaths, which are generally harder to miss, were themselves growing exponentially in late January at around 0.22 to 0.27 per day. Slower than cases, as you'd expect given the lag to death, but marching to the same drumbeat.

There's one more layer here that matters for policy: sensitivity. They pushed on twenty-three different "what if" scenarios to see how assumptions about surveillance change the answer. Suppose a constant fraction of infections was never detected outside Hubei.

That shifts t-zero earlier—you had more infections than you thought—but it barely nudges the growth rate. Suppose surveillance outside Hubei ramped up over January. That can pull the estimated r down.

In their most conservative setup, the growth rate dips to about 0.21 per day. High either way. The point is, under very different detection stories, you still need a rapidly growing outbreak inside Wuhan to explain what the rest of China was seeing.

Now, growth rate is one thing. Transmissibility is another. The bridge between them is the serial interval—the time between when one person gets infected and when the people they infect get infected—which depends on how long people incubate the virus and how long they're infectious.

Using early Wuhan estimates that put the serial interval around seven to eight days, Sanche and colleagues back out a basic reproductive number, R-naught, of about 5.8. The ninety-five percent interval runs from 4.4 to 7.7. Widen that serial-interval range to six to nine days and the median barely moves—to about 5.7—with an uncertainty window from 3.8 to 8.9.

That's uncomfortably high. Much higher than the early, low-twos R-naught values that floated through the first headlines.

Those R-naught values come with a quiet corollary. If each infected person infects, on average, between four and eight others in a fully susceptible population, single interventions won't cut it. You need layers: surveillance to catch cases earlier, quarantine and contact tracing to box in chains of transmission, and strong social distancing to pull the average number of contacts way down.

That's not spin; it's arithmetic. In a classic herd-immunity calculation, the fraction of the population that needs to be immune or insulated from infection is one minus one over R-naught. When that denominator is five or six, the target for control moves far out.

Along the way, they also extract the time constants of the disease from outside Hubei case reports. The average incubation period—the time from exposure to symptoms—lands around 4.2 days, with a plausible range from 3.5 to 5.1. There's a telling shift in behavior as awareness grew: the time from symptom onset to hospitalization dropped from about 5.5 days before January 18 to 1.5 days after, a change that's statistically robust.

The time from symptom onset to death averages about 16.1 days. Those numbers set the pace of the clinical course and explain why case counts and death counts look like offset waves.

Hospital stays track with that story. On average, people discharged recovered after about eleven and a half days in the hospital, while those who died did so after roughly eleven days of hospitalization. That's the human timeline behind the transmission math, and it frames why health systems felt slammed: fast-growing infections layered on multi-week clinical arcs.

One subtle but important assumption underlies both models: exported cases were detected perfectly. That's a strong claim in a chaotic moment. Sensitivity checks help.

If, in reality, a fixed chunk of infected travelers was missed, the models respond by pushing the onset of exponential growth earlier, but the slope of that growth—the daily r—stays about the same. If detection was improving over time, some of what looks like exponential growth in exported case counts is just better surveillance. That's why their worst-case growth rate is lower, at about 0.21 per day. Even then, the doubling time is only about three days. Not exactly reassuring.

It's worth pausing on how these methods read the epidemic through motion. The Baidu data are not lab tests or case forms. They're anonymized traces of where people went.

Tying those flows to when the very first infected travelers surfaced in each province turns mobility into a measurement tool. In parallel, treating Wuhan's internal growth deterministically while letting the exported infections play out stochastically matches the reality that big populations obey smooth curves, while the first few cases in a new place are jumpy and discrete. As a cross-check, the two models don't just agree with themselves; they align with death trajectories and with independent lines of evidence on generation times.

One more uncomfortable detail emerges in their broader checks: perhaps a fifth of transmission could have been coming from unidentified cases. If that's true, contact tracing alone can't chase it all down. It pushes you toward population-level tools—closing gathering places, widespread masking, workplace changes—that don't care whether someone has been named a case yet.

And it tightens the clock. When your doubling time is two to three days, every day you wait multiplies the problem.

Why does any of this history lesson matter now? Because it shows how to read an outbreak even when your central data are smoky. Watch the periphery.

Measure movement. Use simple equations you can explain at a whiteboard. And stress-test your conclusions against the biases you can't avoid.

Sanche and colleagues show that under a wide range of assumptions, the early Wuhan outbreak was on a trajectory that would swamp hospitals unless it was met with early, layered controls. That conclusion isn't just about one city in one month. It's about the physics of fast epidemics.

When you turn the math back into policy, the message is blunt. With growth rates around 0.21 to 0.30 per day and R-naught hovering near six under plausible intervals, you need surveillance that's faster than people's social lives, quarantine and tracing that run on tight clocks, and social distancing strong enough to slash contact rates quickly. Wuhan's intense pre-lockdown mobility—those forty thousand to one hundred forty thousand daily departures—explains how sparks landed across China.

The models explain why dampening those sparks needed to happen everywhere at once.

There are caveats, and they're honest ones. Early case reports tilted toward severe cases. Detection outside Hubei was assumed to be perfect, which it wasn't.

Parameter choices—how long people incubate, how long they're infectious—nudge R-naught up or down. But the central picture holds across the tests they ran: rapid exponential growth, early December takeoff, high transmissibility.

If you want a coda, it's this. The models were simple by design. I-star equals a constant times an exponential.

A clean SEIR engine for exported cases. Baidu's travel counts as a backbone. In a blizzard of uncertainty, simplicity was a feature, not a bug, because it let the team tie timing, movement, and growth into a story you can both believe and act on.

And it set a standard for how to infer what matters most in the early days of a new pathogen: how fast it's growing, how widely it can spread, and how quickly you need to move to beat it.

More in Mathematics