A new coronavirus associated with human respiratory disease in China

Fan Wu, Zhao Su, Bin Yu, Yanmei Chen, Wen Wang, Zhi-Gang Song, Yi Hu, Zhao-Wu Tao, Jun-Hua Tian, Yuan-Yuan Pei, Ming-Li Yuan, Yu-Ling Zhang, Fa-Hui Dai, Yi Liu, Qimin Wang, Jiao-Jiao Zheng, Lin Xu, Edward C. Holmes, Yǒng-Zhèn ZhāngView original
OverviewBalancedalloy voice
Picture a single hospital room in Wuhan during the winter of 2019. A 41-year-old man is struggling to breathe. He has a fever, cough, chest tightness, and the whole list. He's already on antibiotics, antivirals, and even steroids, and he's getting worse. Doctors move him to intensive care and start high-flow oxygen. Standing at his bedside is a very simple question with very high stakes: what on earth is causing this pneumonia? Standard tests came first. Think of the greatest hits of wintertime infections: influenza A and B, adenovirus, and the usual atypical bacteria like Mycoplasma and Chlamydia pneumoniae. All tests returned negative by polymerase chain reaction. So Wu and colleagues did something exacting and agnostic: they took fluid from deep in his lungs through bronchoalveolar lavage and sequenced everything in it. Not a panel of suspects, but everything. This is meta-transcriptomics, and it's the lab equivalent of turning on the lights and seeing what's actually in the room. From that sample, they generated a huge pile of data — about 56.6 million paired-end reads on an Illumina instrument — and assembled them without a reference, using de novo tools like Megahit and Trinity. Here's the moment of truth. Among hundreds of thousands of assembled pieces, one long contig, 30,474 nucleotides long, stood out. When they compared it to known sequences, it looked most like a bat SARS-like coronavirus called SL-CoVZC45, with 89.1 percent nucleotide identity. That's not “same virus” close, but it's close enough to ring alarms. They used that long piece to design primers, filled in the ends with rapid amplification of complementary DNA ends, and closed a full genome. They called it WH-Human 1 coronavirus at the time — WHCV — what the world soon learned as the novel coronavirus. Now, how confident can you be that you've got the right genome and not a jumble of fragments from different bugs? They took the raw reads and mapped them back to the new genome. A hundred twenty-three thousand six hundred thirteen reads landed on target, and coverage across the genome was essentially complete — 99.99 percent. The average depth per base was modest, about six-fold on average, but consistent with a single dominant viral genome in that sample. Layer on top the quantitative PCR estimate of viral load — around three point ninety-five times ten to the eighth copies per milliliter of lung fluid — and you have a picture of an abundant virus driving disease in a specific place in the body. Let's talk anatomy — of the genome. Coronaviruses are modular, and WHCV fit the mold. The complete sequence was 29,903 nucleotides, bracketed by five-prime and three-prime termini of 265 and 229 nucleotides, a pattern you see in betacoronaviruses. The gene order ran from the big replicase, ORF1ab, into the spike protein, then the small envelope, membrane, and nucleocapsid proteins. The predicted ORF1ab spanned a little over 21,000 nucleotides and encoded 16 non-structural proteins — all the enzymatic machinery coronaviruses use to copy themselves. Downstream sat at least 13 additional open reading frames. The spike, envelope, membrane, and nucleocapsid genes were 3,822, 828, 669, and 1,260 nucleotides long, respectively. There was even an ORF8, tucked between membrane and nucleocapsid, that's characteristic of SARS-like viruses. And the transcriptional signals — short motifs coronaviruses use to make their subgenomic messages — were in the right places with the right core sequences. In other words, the architecture looked canonical. Where did this new virus sit on the coronavirus family tree? Wu and colleagues ran phylogenies on the whole genome and on individual genes. No matter how they sliced it, WHCV fell into the subgenus Sarbecovirus, the same branch that includes SARS-CoV and a menagerie of bat SARS-like coronaviruses. But — and this is telling — its precise position shifted depending on the gene you looked at. In the spike gene tree, WHCV's closest neighbor was that bat virus SL-CoVZC45, with about 82 percent amino acid identity in spike and roughly 77 percent to SARS-CoV. In the ORF1b region, which encodes the core polymerase, WHCV dropped into a more basal spot within sarbecoviruses, as if you were looking at a slightly different history. When related viruses refuse to tell a single, clean family story, recombination is usually the narrator in the background. To test that idea directly, the team used the Recombination Detection Program, RDP4, and a companion analysis called SimPlot. Think of these as statistical tools that scan for stretches of sequence that swap their closest relatives, a hallmark of past recombination. Looking across the full genome, there wasn't strong evidence that WHCV was a patchwork. But zoom in on spike, and the picture changes. Two breakpoints popped out, at nucleotides 1,029 and 1,652, carving spike into three segments. The outer segments — from the start to 1,029 and from 1,652 to the end — grouped with bat SARS-like viruses such as ZC45 and ZXC21. The middle piece, from 1,030 to 1,651, which happens to include the receptor-binding domain, clustered with SARS-CoV and bat viruses like WIV1 and RsSHC014 that can enter human cells. The statistics weren't borderline; p-values ranged from about three times ten to the negative three down to nine times ten to the negative nine across different methods. So spike looked mosaic, even if the rest of the genome did not. Why does that central slice of spike matter so much? Because that's the domain that grabs the host receptor. For SARS-CoV, and for several bat SARS-like viruses, that receptor is angiotensin-converting enzyme 2 — ACE2 — on human airway cells. Wu and colleagues wanted to know whether WHCV's receptor-binding domain looked like the kind that could use human ACE2. They started with sequences. Compared to SARS-CoV's receptor-binding domain, WHCV's was roughly three-quarters identical in amino acids — about 74 percent — and it was even closer to several bat viruses known to use human ACE2, like Rs4874, Rs7327, and Rs4231, at around 76 percent. It was only one amino acid longer than SARS-CoV's receptor-binding domain, with a small insertion at a position analogous to valine 470. Contrast that with a bat virus called Rp3, which can't use human ACE2 and carries deletions in exactly the loops that contact the receptor. Then they went three-dimensional, at least in silico. Using SWISS-MODEL with ProMod3, they built homology models of WHCV's receptor-binding domain and compared them to solved structures: the SARS-CoV receptor-binding domain alone and in complex with human ACE2, structures you may have seen labeled 2GHV and 2AJF in the Protein Data Bank. Model quality was checked with the QMEAN score. The predicted shape of WHCV's receptor-binding domain hugged the SARS-CoV template pretty well, especially in the receptor-contacting loops. Rp3's receptor-binding domain, by comparison, drifted. Put those two strands together — receptor-binding domain sequence closely matching ACE2-using relatives, and a modeled structure that overlays with the SARS-CoV receptor-binding domain at the business end — and you get a cautious but concrete inference: this new virus might well use human ACE2. Wu and colleagues were careful. They didn't claim proof. No binding assays here and no pseudovirus entry experiments. But if you were triaging which hypotheses to test first in the middle of an outbreak, this one would move to the top of the pile. All of this work happened fast, and it wasn't meant to be a randomized, blinded trial. It was a discovery effort anchored in a single patient and a single bronchoalveolar sample handled under high biocontainment. As a snapshot, it was powerful. It gave the world a complete genome — the one you now know as SARS-CoV-2 — and the key numbers to trust it. It also provided public identifiers so anyone could download and analyze the data: the genome under GenBank accession MN908947, and the raw reads in the Sequence Read Archive under BioProject PRJNA603194. The team used the name WHCV or 2019-nCoV because that's what existed then; the International Committee on Taxonomy of Viruses later standardized the name to SARS-CoV-2, and the disease became COVID-19. If you're wondering why this genomic anatomy lesson matters outside a sequencing core, think about what it unlocked overnight. Once you know the sequence, you can design diagnostic primers that hit unique regions and miss everything else, which is exactly what followed. You can place the virus in its evolutionary neighborhood, which pointed not to influenza or an entirely new family, but to sarbecoviruses that circulate in bats. You can scan for recombination, not to sensationalize origins, but to understand how the spike gene — the key to host range — has traded pieces among relatives. And even without a single wet-lab receptor assay, you can make a defensible call that human ACE2 is likely the entry door, which focuses cell biology and animal studies on the right pathways. There are limits here, and the authors were explicit about them. One patient isn't an epidemiology study. A genome in hand doesn't, by itself, prove causation of disease. And in recombination analysis, seeing mosaicism in spike doesn't mean recombination made this virus emerge; sarbecoviruses recombine a lot, and those signatures accumulate over long arcs of evolution. But the combination of negative tests for usual suspects, a lung sample teeming with a single coronavirus genome, and clinical severity consistent with viral pneumonia built a strong circumstantial case — and ignited the investigative engine that confirmed it. The technical choices deserve a little respect, too, because they're part of why this worked. An unbiased, meta-transcriptomic readout meant they didn't miss a novel pathogen just because it wasn't on a test panel. De novo assembly, rather than mapping only to known references, let a 30 kilobase genome emerge from scratch. Mapping reads back to the assembly — with more than a hundred thousand piling on and near-complete coverage — gave the confidence you need before you start designing polymerase chain reaction assays for hospitals. And the simple fact that spike's receptor-binding domain is both sequence-similar to ACE2 users and structurally compatible with the ACE2 interface told virologists where to aim their early experiments. Looking back, this study feels like the opening scene of a story we all now know too well. But in that first chapter, the work did three concrete things. It identified a new sarbecovirus in a critically ill patient, with a 29,903-nucleotide genome that looks and behaves like a coronavirus should. It mapped that virus into the family tree and showed that, while the genome as a whole didn't scream recombination, the spike gene carried a mosaic history with a receptor-binding domain aligned to ACE2-using relatives, strengthened by breakpoint statistics. And it put actionable information into the public domain immediately, seeding diagnostics, basic science, and surveillance. If you want to take one forward-looking note from all this, keep it measured. Genomics doesn't replace virology, but in an outbreak, it sets the table. The sequence pointed to ACE2, and weeks later, functional studies would confirm that. The family tree pointed to bats, and over time, sampling would expand the set of close relatives. The cleanup step — ruling out the usual pathogens — mattered just as much as the shiny assembly statistics. And perhaps most important, the speed and openness with which the data were shared — that GenBank accession, that Sequence Read Archive project — are part of why the global response could move from speculation to experiments so quickly.

Picture a single hospital room in Wuhan during the winter of 2019. A 41-year-old man is struggling to breathe. He has a fever, cough, chest tightness, and the whole list.

He's already on antibiotics, antivirals, and even steroids, and he's getting worse. Doctors move him to intensive care and start high-flow oxygen. Standing at his bedside is a very simple question with very high stakes: what on earth is causing this pneumonia?

Standard tests came first. Think of the greatest hits of wintertime infections: influenza A and B, adenovirus, and the usual atypical bacteria like Mycoplasma and Chlamydia pneumoniae. All tests returned negative by polymerase chain reaction.

So Wu and colleagues did something exacting and agnostic: they took fluid from deep in his lungs through bronchoalveolar lavage and sequenced everything in it. Not a panel of suspects, but everything. This is meta-transcriptomics, and it's the lab equivalent of turning on the lights and seeing what's actually in the room.

From that sample, they generated a huge pile of data — about 56.6 million paired-end reads on an Illumina instrument — and assembled them without a reference, using de novo tools like Megahit and Trinity. Here's the moment of truth. Among hundreds of thousands of assembled pieces, one long contig, 30,474 nucleotides long, stood out.

When they compared it to known sequences, it looked most like a bat SARS-like coronavirus called SL-CoVZC45, with 89.1 percent nucleotide identity. That's not “same virus” close, but it's close enough to ring alarms. They used that long piece to design primers, filled in the ends with rapid amplification of complementary DNA ends, and closed a full genome.

They called it WH-Human 1 coronavirus at the time — WHCV — what the world soon learned as the novel coronavirus.

Now, how confident can you be that you've got the right genome and not a jumble of fragments from different bugs? They took the raw reads and mapped them back to the new genome. A hundred twenty-three thousand six hundred thirteen reads landed on target, and coverage across the genome was essentially complete — 99.99 percent.

The average depth per base was modest, about six-fold on average, but consistent with a single dominant viral genome in that sample. Layer on top the quantitative PCR estimate of viral load — around three point ninety-five times ten to the eighth copies per milliliter of lung fluid — and you have a picture of an abundant virus driving disease in a specific place in the body.

Let's talk anatomy — of the genome. Coronaviruses are modular, and WHCV fit the mold. The complete sequence was 29,903 nucleotides, bracketed by five-prime and three-prime termini of 265 and 229 nucleotides, a pattern you see in betacoronaviruses.

The gene order ran from the big replicase, ORF1ab, into the spike protein, then the small envelope, membrane, and nucleocapsid proteins. The predicted ORF1ab spanned a little over 21,000 nucleotides and encoded 16 non-structural proteins — all the enzymatic machinery coronaviruses use to copy themselves. Downstream sat at least 13 additional open reading frames.

The spike, envelope, membrane, and nucleocapsid genes were 3,822, 828, 669, and 1,260 nucleotides long, respectively. There was even an ORF8, tucked between membrane and nucleocapsid, that's characteristic of SARS-like viruses. And the transcriptional signals — short motifs coronaviruses use to make their subgenomic messages — were in the right places with the right core sequences. In other words, the architecture looked canonical.

Where did this new virus sit on the coronavirus family tree? Wu and colleagues ran phylogenies on the whole genome and on individual genes. No matter how they sliced it, WHCV fell into the subgenus Sarbecovirus, the same branch that includes SARS-CoV and a menagerie of bat SARS-like coronaviruses.

But — and this is telling — its precise position shifted depending on the gene you looked at. In the spike gene tree, WHCV's closest neighbor was that bat virus SL-CoVZC45, with about 82 percent amino acid identity in spike and roughly 77 percent to SARS-CoV. In the ORF1b region, which encodes the core polymerase, WHCV dropped into a more basal spot within sarbecoviruses, as if you were looking at a slightly different history.

When related viruses refuse to tell a single, clean family story, recombination is usually the narrator in the background.

To test that idea directly, the team used the Recombination Detection Program, RDP4, and a companion analysis called SimPlot. Think of these as statistical tools that scan for stretches of sequence that swap their closest relatives, a hallmark of past recombination. Looking across the full genome, there wasn't strong evidence that WHCV was a patchwork.

But zoom in on spike, and the picture changes. Two breakpoints popped out, at nucleotides 1,029 and 1,652, carving spike into three segments. The outer segments — from the start to 1,029 and from 1,652 to the end — grouped with bat SARS-like viruses such as ZC45 and ZXC21.

The middle piece, from 1,030 to 1,651, which happens to include the receptor-binding domain, clustered with SARS-CoV and bat viruses like WIV1 and RsSHC014 that can enter human cells. The statistics weren't borderline; p-values ranged from about three times ten to the negative three down to nine times ten to the negative nine across different methods. So spike looked mosaic, even if the rest of the genome did not.

Why does that central slice of spike matter so much? Because that's the domain that grabs the host receptor. For SARS-CoV, and for several bat SARS-like viruses, that receptor is angiotensin-converting enzyme 2 — ACE2 — on human airway cells.

Wu and colleagues wanted to know whether WHCV's receptor-binding domain looked like the kind that could use human ACE2. They started with sequences. Compared to SARS-CoV's receptor-binding domain, WHCV's was roughly three-quarters identical in amino acids — about 74 percent — and it was even closer to several bat viruses known to use human ACE2, like Rs4874, Rs7327, and Rs4231, at around 76 percent.

It was only one amino acid longer than SARS-CoV's receptor-binding domain, with a small insertion at a position analogous to valine 470. Contrast that with a bat virus called Rp3, which can't use human ACE2 and carries deletions in exactly the loops that contact the receptor.

Then they went three-dimensional, at least in silico. Using SWISS-MODEL with ProMod3, they built homology models of WHCV's receptor-binding domain and compared them to solved structures: the SARS-CoV receptor-binding domain alone and in complex with human ACE2, structures you may have seen labeled 2GHV and 2AJF in the Protein Data Bank. Model quality was checked with the QMEAN score.

The predicted shape of WHCV's receptor-binding domain hugged the SARS-CoV template pretty well, especially in the receptor-contacting loops. Rp3's receptor-binding domain, by comparison, drifted. Put those two strands together — receptor-binding domain sequence closely matching ACE2-using relatives, and a modeled structure that overlays with the SARS-CoV receptor-binding domain at the business end — and you get a cautious but concrete inference: this new virus might well use human ACE2.

Wu and colleagues were careful. They didn't claim proof. No binding assays here and no pseudovirus entry experiments.

But if you were triaging which hypotheses to test first in the middle of an outbreak, this one would move to the top of the pile.

All of this work happened fast, and it wasn't meant to be a randomized, blinded trial. It was a discovery effort anchored in a single patient and a single bronchoalveolar sample handled under high biocontainment. As a snapshot, it was powerful.

It gave the world a complete genome — the one you now know as SARS-CoV-2 — and the key numbers to trust it. It also provided public identifiers so anyone could download and analyze the data: the genome under GenBank accession MN908947, and the raw reads in the Sequence Read Archive under BioProject PRJNA603194. The team used the name WHCV or 2019-nCoV because that's what existed then; the International Committee on Taxonomy of Viruses later standardized the name to SARS-CoV-2, and the disease became COVID-19.

If you're wondering why this genomic anatomy lesson matters outside a sequencing core, think about what it unlocked overnight. Once you know the sequence, you can design diagnostic primers that hit unique regions and miss everything else, which is exactly what followed. You can place the virus in its evolutionary neighborhood, which pointed not to influenza or an entirely new family, but to sarbecoviruses that circulate in bats.

You can scan for recombination, not to sensationalize origins, but to understand how the spike gene — the key to host range — has traded pieces among relatives. And even without a single wet-lab receptor assay, you can make a defensible call that human ACE2 is likely the entry door, which focuses cell biology and animal studies on the right pathways.

There are limits here, and the authors were explicit about them. One patient isn't an epidemiology study. A genome in hand doesn't, by itself, prove causation of disease.

And in recombination analysis, seeing mosaicism in spike doesn't mean recombination made this virus emerge; sarbecoviruses recombine a lot, and those signatures accumulate over long arcs of evolution. But the combination of negative tests for usual suspects, a lung sample teeming with a single coronavirus genome, and clinical severity consistent with viral pneumonia built a strong circumstantial case — and ignited the investigative engine that confirmed it.

The technical choices deserve a little respect, too, because they're part of why this worked. An unbiased, meta-transcriptomic readout meant they didn't miss a novel pathogen just because it wasn't on a test panel. De novo assembly, rather than mapping only to known references, let a 30 kilobase genome emerge from scratch.

Mapping reads back to the assembly — with more than a hundred thousand piling on and near-complete coverage — gave the confidence you need before you start designing polymerase chain reaction assays for hospitals. And the simple fact that spike's receptor-binding domain is both sequence-similar to ACE2 users and structurally compatible with the ACE2 interface told virologists where to aim their early experiments.

Looking back, this study feels like the opening scene of a story we all now know too well. But in that first chapter, the work did three concrete things. It identified a new sarbecovirus in a critically ill patient, with a 29,903-nucleotide genome that looks and behaves like a coronavirus should.

It mapped that virus into the family tree and showed that, while the genome as a whole didn't scream recombination, the spike gene carried a mosaic history with a receptor-binding domain aligned to ACE2-using relatives, strengthened by breakpoint statistics. And it put actionable information into the public domain immediately, seeding diagnostics, basic science, and surveillance.

If you want to take one forward-looking note from all this, keep it measured. Genomics doesn't replace virology, but in an outbreak, it sets the table. The sequence pointed to ACE2, and weeks later, functional studies would confirm that.

The family tree pointed to bats, and over time, sampling would expand the set of close relatives. The cleanup step — ruling out the usual pathogens — mattered just as much as the shiny assembly statistics. And perhaps most important, the speed and openness with which the data were shared — that GenBank accession, that Sequence Read Archive project — are part of why the global response could move from speculation to experiments so quickly.

More in Medicine