Incubation period of 2019 novel coronavirus (2019-nCoV) infections among travellers from Wuhan, China, 20–28 January 2020
Picture late January 2020. Airports were full of masks, headlines were full of questions, and there was one deceptively simple clock everyone needed to set: how long from getting infected to feeling sick. That interval, the incubation period, decides who counts as a contact, how long you quarantine, and even whether airport screening has any chance of catching cases.
In the middle of that fog, Backer, Klinkenberg, and Wallinga found an unlikely source of clarity: travel itineraries.
They noticed a natural experiment unfolding. Dozens of people left Wuhan, later tested positive abroad, and many had precise dates for when they were in the city and when their symptoms began. That’s a gift to an epidemiologist.
If you know the window when someone could have been exposed—say, the five days they spent in Wuhan—and you know when symptoms started, then the incubation time is somewhere between "onset minus the earliest exposure" and "onset minus the latest exposure." It’s not a single number per person; it’s an interval. Multiply that by a cohort, and you can shape a distribution. In their dataset, drawn from translated Chinese reports and international confirmations up to January twenty-ninth, there were eighty-eight travelers with this kind of timeline.
Turning itineraries into evidence means embracing uncertainty rather than pretending it away. Backer and colleagues treated each case as interval-censored: the infection happened sometime within the travel window, and the symptom onset was fixed. They then used a Bayesian model to learn the shape of the incubation-period distribution from those intervals.
Two design choices mattered. First, they put a uniform prior over the exposure window for each person, which is a fancy way of saying "infection was equally likely on any day of the stay." Second, they used weak, strictly positive priors on the shape and scale parameters of the distribution, allowing the data to do most of the talking. The computation ran in Stan, via the rstan package in the R programming language, which let them draw from the full posterior and carry the interval-censoring all the way into the credible intervals.
They didn’t bet everything on one mathematical form. Instead, they fit three common candidates to describe time-to-event data: Weibull, gamma, and lognormal. Then they asked a simple question: which one describes these eighty-eight cases best?
Using leave-one-out information criteria—a cross-validation score where lower is better—the Weibull came out on top with a value of 486, the gamma was next at 545, and the lognormal trailed at 592. So the general contour was clear. Still, the key policy question is not which curve wins a horse race. It’s where that curve sits and how fat its tail is.
Here’s the headline. The mean incubation period centered around six days. Under the best-fitting Weibull model, the mean was 6.4 days, with a 95 percent credible interval from 5.6 to 7.7.
Take a breath and map that to real life: if you were exposed on a Monday, symptoms on average would show up on Sunday, but there’s real spread around that average.
The tails tell you about caution. Under the Weibull fit, the central ninety-five percent of cases fell between about 2.1 and 11.1 days after infection. The gamma presented a similar story but allowed a little more room at the top, stretching to roughly 12.5 days.
The lognormal pushed the upper end further, up toward 15.5 days. A different way to see this is to look at the ninety-fifth percentile: about 10.3 days for Weibull, 11.3 for gamma, and around 13.3 for lognormal. The same six-day center, but different degrees of "just in case."
That difference in tails matters for quarantine policies. It explains why a fourteen-day quarantine became the cautious default around the world. Backer and colleagues noted that, even though the lognormal fit the data worst overall, its heavier upper tail kept the possibility of very long incubations on the table.
And yet, the best-fitting curve said most people would declare themselves sooner than two weeks, with only a small fraction lagging toward that boundary.
The team also checked how their conclusions shifted when they focused on the cleanest subset: travelers whose exposure windows were completely closed—meaning you knew exactly when they entered and left Wuhan. In those twenty-five cases, the estimated mean incubation dropped to 4.5 days, with a credible interval from 3.7 to 5.6, and the ninety-fifth percentile came in around eight days. That’s a notable difference.
It indicates that some of the six-day center in the full sample could be nudged upward by fuzzier windows or by an artifact of early outbreak data, where people who got sick soon after travel were easier to identify and report.
Methodologically, there’s a subtlety worth pausing on. When a maximum plausible incubation time was available for a person—because, for example, they left Wuhan on a known date—the inferred infection times tended to cluster toward the end of their stay. That’s intuitive: if you got sick three days after leaving and the incubation could be as short as two days, the model leans toward you being infected near the tail end of your trip.
The danger is that this pattern, repeated across cases, can make incubation periods look a bit longer overall, especially early on when right-truncation and reporting delays filter who shows up in the data.
So, what do we learn if we zoom out from the statistics to the policy levers people were pulling? First, case definitions and contact-tracing windows around a week made sense. Second, border screening that relied on catching fevers within a couple of days of exposure was never going to be highly effective; the lower two point five percent percentile was still around two days.
Third, fourteen days for quarantine was conservative but defensible, particularly if you were worried about the small probability mass out in the lognormal tail.
Now, whenever a new virus shows up, we reach for familiar cousins. How did these numbers line up against SARS and Middle East respiratory syndrome? Surprisingly well.
The mean for 2019-nCoV sat within about a day of typical MERS estimates, which hovered between roughly five and seven days in studies by Assiri, Cauchemez, and Virlogeux. The high end wasn’t dramatically different either; the ninety-fifth percentiles for MERS often fell in the eleven to thirteen day range, and COVID’s Weibull and gamma fits lived right next door. SARS was the messy sibling.
Depending on the study—Donnelly’s early work in 2003, Cowling’s analysis in 2007, or Lau’s later comparisons across Hong Kong, Beijing, and Taiwan—the mean could be as low as about four days or as high as nearly seven, and the ninety-fifth percentile in some settings ran up close to twenty days. Against that backdrop, it wasn’t crazy in January 2020 to borrow from MERS or mid-range SARS priors when building early models for COVID-19. Backer and colleagues effectively gave people permission to do that without feeling reckless.
Let’s circle back to the model comparison because it has a practical punchline. If you’re deciding on policy, you care about two things: central tendency and tail risk. The three distributions agreed on the center—call it six days—so you’re not gambling much there.
They disagreed on how much probability to assign beyond ten, twelve, or fifteen days. The leave-one-out comparison favored the Weibull, and that’s the shape that kept the tail relatively tight. But the team didn’t hide the fact that if you picked a lognormal, which has a reputation for hefty right tails in biological time processes, you’d be planning for a few more late bloomers. That transparency helped public health agencies choose how conservative to be.
What about the caveats? They’re real. This was a traveler-biased sample: more male, more mobile, and potentially younger than the general patient pool.
Exposure histories came from reports, which means recall and reporting bias. The modeling assumed a uniform chance of infection across the stay in Wuhan; if, in reality, people were more likely to be infected on peak exposure days—crowded markets or hospital visits—that would bend the true incubation distribution a bit. And because the data were drawn so early in the outbreak, when only cases that had already declared themselves abroad could be counted, there’s that right-truncation bias again nudging estimates upward.
The authors flagged this clearly and called their upper limit—11.1 days from the two point five percent to ninety-seven point five percent range under the Weibull—conservative.
Still, for what it set out to do—give decision makers numbers they could use in the moment—the study delivered. It offered a mean incubation of 6.4 days with a measured uncertainty band. It drew a line around where most cases would land—about two to eleven days under the best-fitting shape—and mapped how a more cautious choice of distribution could push that boundary out.
It showed that a cleaner subset of exposure windows tilted shorter, hinting at biases people would need to revisit as better data arrived. And it anchored those findings in the wider coronavirus family, easing the adoption of SARS- and MERS-informed policies without overselling the analogy.
There’s also a quiet methodological win here. Interval-censoring is not a glamorous phrase, but it’s the honest way to treat the data you usually have in an outbreak: ranges, not timestamps. By building a Bayesian model that respects those ranges and by sampling the full posterior, Backer and colleagues made every assumption visible—what they believed about infection timing within a stay, how much they trusted the shape parameters, and how sensitive the answer was to those choices.
That kind of transparency is as important as the headline number when you’re deciding who has to stay home and for how long.
If you’re wondering what happened next, the arc is familiar. As the pandemic unfolded, larger datasets and contact-tracing studies refined these early estimates. But the first cut matters.
It set quarantine guidance in motion, aligned reasonably with the World Health Organization’s up to fourteen-day stance and the European Centre for Disease Prevention and Control’s roughly two to twelve day framing, and gave modelers a backbone for forecasting spread.
One last thought, looking forward. The trick Backer and colleagues pulled—mining travel histories for timing information and fitting flexible time-to-event models—should be in the standard playbook for the next emerging pathogen. Pair it with richer exposure diaries or early serology, and you can narrow those tails faster.
And when the stakes are high and the data are thin, that’s the difference between a policy that’s roughly right and one that’s precisely wrong.
Related lectures
- Estimating the infection and case fatality ratio for coronavirus disease (COVID-19) using age-adjusted data from the outbreak on the Diamond Princess cruise ship, February 2020
- Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
- Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School