Estimating the asymptomatic proportion of coronavirus disease 2019 (COVID-19) cases on board the Diamond Princess cruise ship, Yokohama, Japan, 2020
Picture a cruise ship in winter, sitting in Yokohama Bay. A floating city with three thousand seven hundred eleven people on board, sealed off for two weeks starting February fifth, 2020. It's an eerie scene, but also a scientific gift: a closed population, fixed in place, with a virus spreading inside.
By February twentieth, testing had found six hundred thirty-four infections across three thousand sixty-three polymerase chain reaction tests. One ship, one setting, a tight window of time. As Mizumoto, Kagaya, Zarebski, and Chowell realized, that's an unusually clean way to ask a messy question: how many people infected with this new coronavirus never develop symptoms at all?
That sounds simple. It's not. At the moment someone is swabbed, they might feel fine because they won't ever get sick, or because they're infected and symptoms just haven't arrived yet.
On the Diamond Princess, three hundred six infected people were symptomatic at testing, three hundred twenty-eight felt fine. If you stopped there, you'd be tempted to call it roughly half and half. But the fraction of people asymptomatic at testing climbed over the quarantine, from about sixteen percent before February thirteenth to just over fifty percent by February nineteenth.
That rise wasn't the virus changing; it was the clock. Many of those asymptomatic passengers were simply presymptomatic when the sample was taken.
So the team built the analysis around time. Conceptually, think of each person's infection happening at some unknown moment within a window when they could have been exposed — call that window from a to b. After infection, there's a delay until symptoms appear, if they appear at all.
That incubation delay is well described, at least early in 2020, by a Weibull distribution — a flexible, skewed curve — with a mean of six point four days and a standard deviation of two point three days. The key quantity they wanted, p, is the probability an infected person never develops symptoms. At the censoring time — here, mid-quarantine, treated as February eighteenth — each person is observed as symptomatic or not yet symptomatic.
The math then asks: given a possible infection time X within that window, what's the chance someone appears asymptomatic at the observation time? It's the sum of two paths. Either they're truly asymptomatic (with probability p), or they're destined to develop symptoms but haven't yet by that date (with probability one minus p, multiplied by the chance the incubation delay is longer than the time elapsed since infection).
That person-level probability becomes a likelihood term, and you multiply those terms across everyone on board.
Because infection times were only known to fall within intervals, not exact days, the model integrates over each person's plausible window. And because some people who looked fine on the observation date would develop symptoms later, the data are right-censored, like survival data where not every event has happened by the time you stop the clock. The authors wrapped that structure in a Bayesian framework and used Hamiltonian Monte Carlo — an efficient way to explore the probabilities of many unknowns at once — to estimate both the latent infection times and that target number p.
It's a mouthful to describe, but the idea is clean: connect the biology of incubation to the timing of observation so asymptomatic today doesn't get mistaken for asymptomatic forever.
When they did that, the picture sharpened. The estimated proportion of all infections that were truly asymptomatic — people who would never go on to develop symptoms — was seventeen point nine percent, with a ninety-five percent credible interval from fifteen point five to twenty point two percent. That's the headline.
Not half. Not even a third. Around one in six.
And among those who tested positive but had no symptoms at the time, only about a third were truly asymptomatic; the posterior median there was zero point thirty-five, with a tight interval around it. Translating those fractions back onto the ship, the model points to roughly one hundred thirteen truly asymptomatic infections out of the six hundred thirty-four detected — again, with a reasonably narrow uncertainty band.
Take a breath and consider what that means. On a ship where the raw, in-the-moment counts could mislead you into thinking asymptomatic infections were rampant, a delay-aware analysis says: meaningful, yes; dominant, no. Much of what looked asymptomatic at testing was the virus still ticking toward symptom onset.
The authors didn't leave their conclusions hanging on one incubation assumption. They nudged the mean incubation from five point five to nine point five days to see how sensitive the answers would be. The big message held.
With shorter incubations, fewer people would be misclassified as asymptomatic at the observation time; with longer ones, more would be. Across that sweep, the overall asymptomatic share among all infections ranged from about twenty-one to forty percent. The exact slice among those who felt fine at testing moved, too, from roughly zero point twenty-eight up to zero point forty.
The total count of truly asymptomatic cases tracked those shifts, landing between about ninety-two and one hundred thirty-one. The point isn't to memorize the bounds. It's to see that, under reasonable incubation scenarios, the asymptomatic footprint is substantial but not the majority story on this ship.
Timing didn't just matter for symptoms; it mattered for when people got infected in the first place. Using heat map style reconstructions of individual infection windows, the analysis showed that most infections happened before or right around the start of the quarantine on February fifth. Symptomatic cases bunched near that line; infections that ultimately proved asymptomatic tended to occur even earlier.
That pattern tells you the virus had already seeded itself widely in the days leading up to the quarantine. It's a reminder of how quickly transmission can stack up in a confined environment when people share air, meals, and corridors.
All of this sits on a data backbone that was unusually detailed for such an early outbreak, but still imperfect in ways that matter. Testing on the Diamond Princess wasn't random. In the early days, labs focused on people with symptoms or those at higher risk.
By February twenty-first, the official tally of infections was still six hundred thirty-four, among three thousand seven hundred eleven passengers and crew, but who got tested, and when, shaped what those daily lines looked like. The snapshot the model worked with marked people as symptomatic or not as of February eighteenth, but some cases were inevitably still in motion. There's also the age question.
The passenger list skewed older — among the six hundred thirty-four cases, four hundred seventy-six were sixty or older — and older adults are more likely to notice and report symptoms. That tilts the observed mix toward symptomatic illness and could push the true asymptomatic proportion downward if you don't adjust for it. The authors say as much: an age-standardized analysis would be better, and richer clinical histories would help sort comorbidities from COVID-19 symptoms.
Still, in the context of a real-time public health emergency, this is as careful as it gets. The team treated the observations for what they were — survival-like data with right censoring and interval-censored exposures — and matched them to the biology of symptom timing. They didn't conflate no cough today with will never cough.
They used a Weibull incubation because it captures the right skew we see clinically: a lot of people developing symptoms near the mean, with a tail that stretches later. And they leaned on Hamiltonian Monte Carlo not as a buzzword, but because to credibly quantify uncertainty around both infection times and asymptomatic probabilities, you need a method that can explore a complex landscape of possibilities without getting stuck.
It's also worth sitting with the contrast between the ship's daily updates and the delay-adjusted reality. On the ground — or rather, on the water — officials watched the percentage of asymptomatic positives rise day after day, peaking near half of all new positives by February nineteenth. That was a stressful signal, hinting at silent spreaders everywhere.
The model reframed that signal: a big chunk of those silent cases were presymptomatic, destined to become obvious a few days later. In other words, the data were right, but the interpretation needed a clock.
Now, a quick word about generalizability. The Diamond Princess isn't a perfect microcosm of the world. The people on board were older, not randomly sampled, and living in a very particular environment.
The twenty-eight nationalities represented didn't add up to a demographically balanced cohort. The labs prioritized symptomatic testing early, which risks missing quiet infections. All of that could bias the asymptomatic share downward if older adults are more likely to notice and report symptoms, or upward if early testing skipped mild cases.
Mizumoto and colleagues are clear about these caveats. Their result — about one in six infections truly asymptomatic under their main assumptions, with plausible swings under different incubation clocks — is best read as an anchored estimate in a very specific setting, not a universal constant.
But the methodological lesson travels. If you want to understand the hidden portion of an epidemic, you can't just count who's coughing on the day you sample. You have to model the path from infection to symptoms and respect the fact that, when you stop the tape, many stories aren't finished.
That's what this analysis did. It took a high-profile outbreak that could have produced all kinds of misleading headlines and asked the harder question: given what we know about timing, what fraction of infections truly stay silent?
As a final reflection, think back to the ship in the harbor. Most infections happened before the doors shut, and many of the asymptomatic swabs were simply caught mid-journey. The delay-adjusted lens pushed the apparent asymptomatic tide back to a more modest but still important slice.
In a pandemic's early fog, that kind of clarity matters. It changes how you weigh screening strategies, how you interpret rising counts, and how you talk to the public about risk. And it reminds us that in outbreak science, time isn't a nuisance. Time is the data.
Related lectures
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School
- Early dynamics of transmission and control of COVID-19: a mathematical modelling study
- Serial Interval of COVID-19 among Publicly Reported Confirmed Cases
- First cases of coronavirus disease 2019 (COVID-19) in the WHO European Region, 24 January to 21 February 2020