The Epidemiological Characteristics of an Outbreak of 2019 Novel Coronavirus Diseases (COVID-19) — China, 2020
Picture the last days of 2019 in Wuhan: clusters of pneumonia with no known cause, patients and clinicians worried, and a mystery that felt uncomfortably familiar. Within a week, investigators had the culprit — a new coronavirus — and by mid-January it was clear this virus could spread from person to person. That changed the stakes overnight.
China tightened its public health authority on January 20, then took the extraordinary step of restricting travel in and out of Wuhan on January 23. The question wasn’t just what the virus was. It was how fast it was moving, who it was hitting, and what that meant for a country trying to get ahead of it.
To answer that, China’s Centers for Disease Control and Prevention, or China CDC, did something ambitious and, for that moment, essential. They pulled every COVID-19 case reported nationwide into one dataset, de-identified it, and tried to take a snapshot of the epidemic as of February 11, 2020. The Novel Coronavirus Pneumonia Emergency Response Epidemiology Team relied on a centralized reporting system where each case carried a unique national identification number, which meant no duplicates and no sampling — just a census of what had been diagnosed.
Cases fell into four categories: confirmed, suspected, clinically diagnosed — used only in Hubei when testing lagged — and asymptomatic. They also tagged severity: mild for non-pneumonia or mild pneumonia, severe for fast breathing, low oxygen, or extensive lung infiltrates, and critical for respiratory failure, septic shock, or multiple organ failure.
There are a few plumbing details that matter because they shape the curves you’ll hear about. "Date of onset" wasn’t the lab result; it was the day a patient said their fever or cough began. Exposure to Wuhan was defined plainly — recent residence, a visit there, or close contact with someone infected — and it turned out to be a powerful variable for explaining spread. Not every field in the system was mandatory, so exposure history, comorbidities, and even severity were sometimes missing.
And because testing capacity was tight in those first weeks, a large share of the cases in the database were not yet confirmed by a nucleic acid test.
Now the scale. The team analyzed seventy-two thousand three hundred fourteen unique records through February 11. Of those, forty-four thousand six hundred seventy-two were confirmed infections, sixteen thousand one hundred eighty-six were suspected, ten thousand five hundred sixty-seven were clinically diagnosed in Hubei, and eight hundred eighty-nine were asymptomatic.
That’s the national canvas. When you look just at the confirmed group — the forty-four thousand six hundred seventy-two — you see a center of gravity that will surprise no one who remembers those days. Almost three quarters, seventy-four point seven percent, were diagnosed in Hubei, and among all confirmed cases, eighty-five point eight percent reported a Wuhan-related exposure.
The age profile sat mostly between thirty and seventy-nine years, and the sex split was close to even. The striking thing wasn’t who was getting infected. It was how those infections played out across severity and geography.
Most people in the confirmed group had what the team called mild disease — about four in five, eighty point nine percent. Severe illness accounted for thirteen point eight percent, and four point seven percent were critical. Here’s the sobering edge of that distribution: deaths in this dataset occurred in the critical category, and among those critical patients, roughly half died — a case fatality of forty-nine percent.
Step back to the whole confirmed cohort and the headline number settles at two point three percent — one thousand twenty-three deaths out of forty-four thousand six hundred seventy-two. But that average hides a brutal contrast. In Hubei, the confirmed case fatality rate was two point nine percent; outside Hubei it was zero point four percent.
Same virus, same country, different context — and the context was a health system under extraordinary strain.
You can see that strain in time as well as space. The team plotted epidemic curves two ways: by the day symptoms started and by the day a laboratory confirmed the diagnosis. Nationally, symptom onset for confirmed cases clustered in late January and peaked around February 1.
Lab confirmations, though, peaked later — on February 4 — showing the delay you’d expect when tests are scarce and backlogs grow. If you zoom outside Hubei, the onset peak appears earlier, around January 27. That tells a story of seeding from the epicenter followed by a wave that crested quickly in the rest of the country, then ebbed as control measures took hold.
Spread wasn’t just fast. It was wide. By January 20, cases had appeared in six hundred twenty-seven counties across thirty provinces, autonomous regions, and municipalities.
By February 11, the virus had touched one thousand three hundred eighty-six counties across all thirty-one provincial-level regions. That is an extraordinary geographic diffusion in about three weeks, and it maps cleanly onto the period when travel restrictions, quarantines, and a nationwide mobilization of health workers rolled into place.
There’s another metric tucked into the paper that’s wonkier but useful: person-time. Across the confirmed cases, the team counted six hundred sixty-one thousand six hundred nine person-days of observation and calculated a mortality rate of zero point zero one five deaths per ten person-days. Think of it as a speedometer for fatal outcomes during follow-up, rather than a simple proportion of deaths.
It complements the case fatality rate by accounting for the time patients were actually observed.
Risk wasn’t uniform across age or health status, and here the gradients are steep. Among confirmed cases, the fatality rate hit fourteen point eight percent for people eighty and older. For those in their sixties it was three point six percent.
Men saw higher fatality than women — two point eight percent versus one point seven percent — a pattern we saw echoed later in many countries. Comorbidities pulled that risk line up sharply. Cardiovascular disease carried a ten point five percent fatality, diabetes seven point three percent, and hypertension six point zero percent.
If you had no recorded comorbidity, the fatality rate was zero point nine percent. These numbers aren’t destiny, and they reflect just the early period of the outbreak, but they draw a clear map of vulnerability.
Hospitals themselves were a front line for infection. The team counted three thousand nineteen health workers infected across four hundred twenty-two medical facilities, of whom one thousand seven hundred sixteen were confirmed cases, and five died. Two-thirds of those confirmed infections — sixty-four percent — were in Wuhan.
Early on, the proportion of health worker cases that were severe or critical was high. In Wuhan, that figure started at thirty-eight point nine percent in the first third of January. By the first third of February it had fallen to twelve point seven percent.
Nationally, the severe or critical share among health workers dropped from forty-five point zero percent to eight point seven percent over the same windows. Outside Hubei, among two hundred fourteen confirmed health worker cases, seven percent were severe or critical and there were no deaths. That arc — from crisis to better outcomes — speaks to how quickly clinical practice, personal protective equipment, and triage improved once the danger was recognized.
A couple of focused looks sharpen that picture of spread. When the team followed confirmed cases diagnosed outside Hubei, most still traced back to Wuhan exposure — again, eighty-five point eight percent. That’s a classic signature of a single epicenter seeding multiple regions before local transmission takes over.
And when they overlaid onset curves with diagnosis curves, the lag made visible how epidemiology gets distorted when laboratory capacity is playing catch-up. Symptoms crest, but the confirmed case count keeps rising for a few more days because the swabs and machines are working through a queue.
If you’re asking what kind of outbreak this was in those first six weeks, the shapes in the data give a straightforward answer. December looked like a continuous common-source pattern — many people exposed to the same source. January and into early February look like a propagated outbreak — person-to-person transmission with successive waves that grow, crest, and then decline.
That distinction matters because it changes how you deploy the tools we had in early 2020: isolate cases, trace contacts, guard hospitals, and, when transmission outpaces those measures, reduce mobility sharply.
Before we turn this into a neat story about control, the team is careful — and we should be too — about what these numbers can and can’t tell you. Case definitions evolved, especially with the clinically diagnosed category used only in Hubei when laboratory slots were scarce. About thirty-seven percent of all cases in the database were not confirmed by a nucleic acid test at the time.
Several variables, including comorbidities, exposure history, and severity, were missing in a notable fraction of records. And the clock was ticking. Right-censoring — the idea that some people who were alive on February 11 would die later — tends to push the early case fatality rate down.
All of that tilts the picture in ways we can’t fully correct for without later data.
Still, the signal rises above the noise. In about thirty days, a virus identified on January 7 moved from a city outbreak to a national emergency. Most cases were mild, but risk climbed steeply with age and chronic illness.
Hubei faced a much higher fatality rate than the rest of the country, a gap that tracks with an overwhelmed health system and the timing of interventions. Health workers absorbed a heavy early burden, then saw severity fall as protections and practice improved. And the geographic maps look like you’d expect after a holiday travel season collided with a novel pathogen.
Why does an early, imperfect snapshot like this matter now? Because it taught something fundamental about how to read an epidemic in motion. You take onset dates seriously, not just lab confirmations.
You stratify by place, because context changes outcomes. And you protect hospitals, because when they become transmission hubs, everything else gets harder. The China CDC team’s nationwide roll-up did more than count cases; it showed where the fire was hottest, who was at greatest risk of being burned, and which lines of defense were starting to hold.
We’ve learned a lot since February 2020 about variants, vaccines, and long-term outcomes. But that first panoramic view — seventy-two thousand three hundred fourteen cases organized by severity, place, and time — is still instructive. It’s the template for how to get from confusion to clarity in the opening act of an outbreak: build the surveillance pipeline, accept the blind spots you can’t yet fix, and use the clearest signals to aim your next moves.
Related lectures
- Estimating the infection and case fatality ratio for coronavirus disease (COVID-19) using age-adjusted data from the outbreak on the Diamond Princess cruise ship, February 2020
- Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
- Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School