Estimating the infection and case fatality ratio for coronavirus disease (COVID-19) using age-adjusted data from the outbreak on the Diamond Princess cruise ship, February 2020
If you want to know how deadly a new virus is, you count the deaths and divide by the cases. But if people are still dying, that fraction is wrong. The denominator includes patients who haven't yet had a chance to die, so the ratio comes out too low. Adjust too slowly and you underestimate the danger. Study a population that's older than average and you overestimate it for everyone else. In February 2020, a cruise ship trapped in Yokohama harbor accidentally became the cleanest natural experiment available for threading all three of those needles at once. The ship was the Diamond Princess, and the chain of events that turned it into an epidemiological dataset is worth knowing. A passenger with symptom onset on January 19th had disembarked on January 25th and tested positive for COVID-19 in Hong Kong on February 1st. The ship returned to Yokohama on February 3rd and was held in quarantine. Over the following weeks, three thousand sixty-three PCR tests were performed across a closed population of three thousand seven hundred eleven passengers and crew. By February 20th, six hundred nineteen confirmed cases had been identified, which is about seventeen percent of everyone on board. Three hundred eighteen of those were asymptomatic at the time of their positive test. That last number is the key. Because testing was so extensive and the population was sealed, researchers could see infections that routine surveillance would have missed entirely.
Timothy Russell and colleagues at the London School of Hygiene and Tropical Medicine recognized what this situation offered. Their goal was to estimate two related but distinct quantities. The case fatality ratio, often called the CFR, is deaths divided by confirmed cases. The infection fatality ratio, or the IFR, is deaths divided by all infected people, including the asymptomatic ones you'd normally never find. The Diamond Princess gave them a shot at both. But first, they had to solve the delay problem. Here is the core methodological challenge. At any point during an active outbreak, many confirmed cases are still ill. Their outcomes, whether death or recovery, are unknown. If you divide deaths to date by cases to date, you're dividing by a number that's too large, and you get a ratio that's too small. The naive CFR, as Russell and colleagues call it, is systematically biased downward. The fix is to estimate how many of the already-reported cases have had enough time to reach an outcome and use that smaller, corrected number as your denominator instead.
To do that, Russell and colleagues borrowed a lognormal distribution, which is a bell-curve-shaped probability model, fitted by Linton and colleagues to Wuhan hospitalization to death data. That distribution had a mean of thirteen days and a standard deviation of twelve point seven days. Using a distribution rather than a single fixed number means you're asking: for a case confirmed today, what's the probability it would result in death after five days? After thirteen? After twenty-five? Aggregating those probabilities over the entire time series of confirmed cases tells you how many cases should, by now, have resolved. That becomes your corrected denominator, and the result is what the team calls the corrected CFR or cCFR. For a sensitivity check, they also ran the analysis with a raw, non-truncated version of the delay distribution, which had a shorter mean of eight point six days. Then they went one step further. Because the Diamond Princess had identified so many asymptomatic cases, Russell and colleagues could translate their corrected CFR into a corrected IFR, which is the fatality ratio across all infections, not just confirmed ones. They used the measured proportion of asymptomatic versus symptomatic cases on the ship to scale between the two.
The results for the ship itself showed an all-age corrected IFR of one point three percent, with a ninety-five percent confidence interval spanning zero point thirty-eight to three point six. The corrected CFR came out at two point six percent, with a confidence interval from zero point eighty-nine to six point seven. Those are wide ranges — genuinely wide — and that uncertainty is real information. For passengers aged seventy and older, the numbers were considerably higher: a corrected IFR of six point four percent and a corrected CFR of thirteen percent, reflecting the steep age gradient that COVID-19 would become known for. But here's the problem with stopping there. The Diamond Princess was not a representative population. Its mean passenger age was fifty-eight, and all deaths on board occurred in people aged seventy or older. Applying those all-age estimates directly to a younger national population would overstate the risk substantially. So Russell and colleagues ran an age standardization. They took age-stratified naive CFR estimates from Chinese surveillance data and applied them to the Diamond Princess age distribution, asking, in effect, how many deaths would you expect on this ship if the Chinese age-specific rates were correct?
The answer was fifteen point fifteen expected deaths. Dividing that by the three hundred one symptomatic cases in the analysis gives an expected naive CFR of five percent for the ship's cohort. The team had observed an actual corrected CFR of two point six percent. Those two numbers, five percent expected versus two point six percent observed, imply that the Chinese naive CFR estimates needed to be scaled down by about fifty-two percent. Apply that scaling to China's overall figures and you get a corrected CFR for China of one point two percent, with a ninety-five percent confidence interval from zero point three to two point seven, and an implied IFR of zero point six percent, ranging from zero point two to one point three. That is six deaths per thousand infections. For comparison, a naive CFR calculated directly from Chinese case and death counts as of early March 2020, specifically two thousand nine hundred eighty-four deaths divided by eighty thousand four hundred twenty-two cases, equaled three point seventy-one percent. The adjustment cuts that headline figure roughly in half, and the IFR cuts it by a factor of six.
That revision mattered enormously at the time. Early alarming figures drawn from reported cases were circulating widely, and they were inflated by exactly the biases Russell and colleagues set out to correct: unresolved outcomes in active cases and a testing pool skewed toward symptomatic, often older patients. The Diamond Princess, by offering near-total infection ascertainment in a closed population, let the team peel those biases apart and quantify them separately. The authors are clear about what their estimates cannot do, and the limitations are worth considering. The ship's passengers may differ from national populations in baseline health, economic status, and access to care on board. Testing prioritized older passengers first, which shaped who got counted early. Healthcare capacity, including how many intensive care unit beds a country has and how quickly patients reach them, influences who survives. This varies across settings in ways this analysis cannot capture. The confidence intervals are wide precisely because the dataset, while unusually clean, was still small: a few hundred confirmed cases, with six deaths on the ship by the time of the analysis.
The broader lesson that Russell and colleagues emphasize is methodological. Real-time fatality estimation is not just arithmetic. It requires modeling the delay from diagnosis to death, accounting for who gets tested and at what ages, and distinguishing between the fraction of confirmed cases who die versus the fraction of all infected people who die. Get any of those wrong and your headline number is misleading — either dangerously low if you ignore outcome delays, or dangerously high if you mistake an elderly cruise ship cohort for the general public. The Diamond Princess didn't give researchers a perfect answer. What it gave them was a rare opportunity to see the virus operating in a controlled, extensively tested population, and to use that window to calibrate the estimates coming out of much larger, noisier surveillance systems. An IFR of around zero point six percent for China, adjusted for age and corrected for delay — that's where careful epidemiology, pressed hard against an accidental dataset, landed in February 2020. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
- Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School
- Early dynamics of transmission and control of COVID-19: a mathematical modelling study