Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
Eleven million people live in Wuhan. In early January 2020, the official case count for a new respiratory illness was forty-one. Those two numbers don't fit. Either this virus was nearly impossible to catch from another person, or health authorities were seeing only a small fraction of what was actually circulating. The entire question facing Wu, Hao, Lau, and their colleagues at the University of Hong Kong was which of those two worlds they were living in — and they had to answer it in real time, with incomplete data, while the clock ran. The outbreak began taking shape in the final days of December 2019. On the thirty-first of December, Wuhan's Municipal Health Commission announced a cluster of viral pneumonia of unexplained origin, and early attention fell on the Southern China Seafood Wholesale Market — a sprawling fifty thousand square metre complex selling seafood, fresh meat, produce, and live wild animals. The earliest known illness onset in the cluster dated back to the eighth of December. The market was disinfected starting on the thirtieth of December and closed on the first of January. By the fifth of January, more than one hundred and sixty close contacts were under medical surveillance, with none found infected. By the eleventh, that number had grown to more than seven hundred people monitored, more than half of them healthcare workers, and still no confirmed secondary infections among them.
A novel coronavirus, subsequently called two thousand nineteen novel coronavirus, was officially identified as the cause on the ninth of January. Of the first forty-one laboratory-confirmed cases, roughly seventy percent reported exposure to the seafood market. But signals of wider spread were already appearing. Thailand reported an imported case with onset on the fifth of January. Japan confirmed one with onset on the third. Some of these exported patients had not visited the Wuhan market at all — they had other wet-market or hospital contacts. Then, around the fifteenth and sixteenth of January, two family clusters surfaced in Guangdong: household members who had never been to Wuhan became infected after relatives returned from the city. By the twentieth of January, fifteen healthcare workers in Wuhan were reported infected. That same day, the Municipal Health Commission reported one hundred and thirty-six new infections in a single day. By the twenty-second of January, the count in publicly available reports had reached four hundred and forty confirmed infections across thirteen provinces and five other countries and regions.
To make sense of all this, Wu and colleagues structured their analysis around two explicit competing scenarios. Scenario one: a large zoonotic spillover — meaning virus jumping from animals to humans — beginning in early December, followed by very limited human-to-human transmission. Scenario two: a smaller initial spillover, then efficient sustained spread from person to person. These weren't just narrative frames. Each scenario made concrete, testable predictions about what the data should look like. Under scenario one, you would expect most cases to have direct market exposure, household and healthcare clusters to be rare, and exported cases to mainly represent travellers who had been near animals. The early case counts supported this reading to a point. If you assumed the market was the sole zoonotic source, the team calculated a basic reproduction number — R naught, the average number of people one infected person passes the virus to in a fully susceptible population — of just 0.3, with a confidence interval running from roughly 0.17 to 0.44. An R naught below one means the outbreak dies out on its own. That would be reassuring news. But scenario two kept forcing its way into the picture. At least three of the exported cases had no contact with the seafood market. Family transmission in Guangdong pointed to person-to-person spread, with one cluster showing a plausible serial interval — the time between successive cases in a transmission chain — of about five days.
Fifteen infected healthcare workers suggested either a super-spreading event or ongoing human-to-human transmission. And if even just one of the first forty-one cases had been infected by another person rather than by an animal, the implied R naught would be 0.02 — one divided by forty-one. A small number, but that calculation illustrates how sensitive these early estimates are to assumptions about the source. The rapid rise from forty-one confirmed cases on the eleventh of January to four hundred and forty by the twenty-second was harder to explain as pure animal spillover. By the twenty-second of January, Wu and colleagues judged the balance of evidence as tilting toward sustained human-to-human transmission — though they stopped short of declaring it confirmed. Then there was the question of severity. How deadly was this virus? The number the world wanted was a case fatality rate. But dividing deaths by total reported cases in a fast-moving outbreak is almost guaranteed to give you the wrong answer. Cases pile up faster than they resolve, so your denominator is artificially large. The team used a different approach: they looked only at resolved hospital cases — patients who had either died or recovered — and calculated the fatality risk as deaths divided by deaths plus recoveries. Only known outcomes counted.
Using the public update available on the twenty-first of January — four deaths and twenty-five recoveries — the formula produced an estimated hospital fatality risk of fourteen percent. The ninety-five percent confidence interval, calculated from the binomial distribution for four deaths out of twenty-nine resolved cases, ran from 3.9 percent to thirty-two percent. That estimate held relatively stable over the ten days since the first death was announced on the eleventh of January. That fourteen percent needs to be understood and questioned simultaneously. It applies only to hospitalized, laboratory-confirmed cases — the sickest patients, the ones who were admitted and whose outcomes were eventually recorded. The paper was direct about what it might not represent. If large numbers of mild or moderate infections were circulating undetected in the community, the true infection fatality rate across all infections could be far lower. The authors suggested it might fall below one percent — potentially even below 0.1 percent — if mild cases were being missed at scale. They also warned that long delays from hospitalization to death in ultimately fatal cases could bias early estimates, and that if new deaths continued to be recorded without corresponding recoveries, the resolved-case formula would itself tend to overestimate risk.
This is the central epistemic problem the paper grapples with: almost every key quantity depends on how many infections are actually out there. The true denominator. If you're only seeing the hospitalizations — the severe tip of the distribution — then your fatality estimate looks high, your reproduction number looks manageable, and your contact tracing looks thorough. Change the denominator and all three shift. The authors named the specific information they needed to resolve this. First: the exposure profile of recently confirmed cases, particularly whether the pattern of non-market-linked infections was growing. Second: case identification and laboratory testing protocols across Wuhan and other cities, since incomplete surveillance could artificially support scenario two by masking many mild cases. Third: the animal reservoir and any intermediary hosts, including supply chains of wild or game meat, which could explain the initial spillover and inform containment. This paper was dated the twenty-second of January, 2020. Within days, the World Health Organization would declare a public health emergency of international concern. The authors did not know that yet.
What they had was a novel pathogen, a few hundred confirmed cases, fragmentary data on clusters and exports, four deaths, and twenty-five recoveries. From that, they constructed a framework clear enough to distinguish between two epidemiologically opposite scenarios, rigorous enough to produce confidence intervals, and honest enough to say which questions it couldn't yet answer. The value wasn't that every number turned out to be right. It was that the reasoning was sound — transparent enough that anyone reading it could see exactly where the uncertainty lived, and why it mattered. When the data is thin and the stakes are enormous, that kind of structured honesty is not a disclaimer. It is the work. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Estimating the infection and case fatality ratio for coronavirus disease (COVID-19) using age-adjusted data from the outbreak on the Diamond Princess cruise ship, February 2020
- Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School
- Early dynamics of transmission and control of COVID-19: a mathematical modelling study