The Diploid Genome Sequence of an Individual Human

Samuel Lévy, Granger Sutton, Pauline C. Ng, Lars Feuk, Aaron L. Halpern, Brian P. Walenz, Nelson Axelrod, Jiaqi Huang, Ewen F. Kirkness, Gennady Denisov, Yuan Lin, Jeffrey R. MacDonald, Andy Wing Chun Pang, Mary Shago, Timothy B. Stockwell, Alexia Tsiamouri, Vineet Bafna, Vikas Bansal, Saul Kravitz, Dana Busam, Karen Beeson, Tina C. McIntosh, Karin Remington, Josep F. Abril, John Gill, Jon Borman, Yu-Hui Rogers, M.E. Frazier, Stephen W. Scherer, Robert L. Strausberg, J. Craig VenterView original
OverviewBalancedjames voice
If every human cell contains two copies of every chromosome — one from your mother and one from your father — then any genome sequence that treats you as a single creature with one copy is missing half the picture. If it's missing half the picture, every variant it finds is an undercount, and every variant it misses could be the one that matters for your health. That's the problem Levy and colleagues set out to fix in two thousand seven, when they sequenced and assembled both copies of one person's genome at once for the first time. That person was J. Craig Venter. What they found was more than four million differences from the reference genome, and nearly three-quarters of the variant bases were not the kind anyone had been systematically counting. Here's what had been missing. The Human Genome Project's reference assembly is a composite drawn from multiple donors, and the earlier Celera assembly is a consensus from five individuals. Both approaches collapse two chromosomal copies into one linear sequence, resulting in a mosaic — a patchwork that doesn't exist on any single chromosome in any real person. Single nucleotide polymorphisms, or SNPs, were the main thing these approaches tracked: positions where one letter in the sequence differs. However, a genome is not just a string of single-letter edits. It contains insertions, deletions, duplications, and inversions — large-scale changes that SNP surveys largely ignore. To see all of that, you need to sequence a real diploid genome and keep the two copies separate. The HuRef project did exactly that. Levy and colleagues generated roughly 32 million sequence reads using Sanger dideoxy technology, with paired-end libraries producing read pairs of about 800 bases from each clone end, achieving approximately 7.5-fold coverage across the genome. The core technical challenge was assembly. Standard assemblers collapse reads from both chromosomal copies into a single consensus, which is precisely what the team wanted to avoid. So they built a modified version of the Celera Assembler that treated each region of variation as a block bounded by nonvariant columns on either side. Reads spanning a variable block were sorted into allele groups; an allele needed at least two supporting reads to be confirmed, and the highest-scoring confirmed allele became the primary consensus while alternate alleles were preserved separately. Think of it as reassembling a shredded book where every page exists in two slightly different editions — the assembler reconstructs both editions rather than averaging them into one that matches neither. The result was 4,528 scaffolds containing 2,810 megabases of contiguous sequence, with a scaffold N50 of 19.5 megabases and roughly 68 percent fewer internal gaps than the earlier whole-genome shotgun assembly. Now for what that assembly revealed. The total variant catalog came to 4,118,889 filtered events, spanning more than 12 million bases. Of those, 3,213,401 are SNPs — about 78 percent of all events. But here is the finding that changes how you think about human genetic variation: the remaining 22 percent of events, the non-SNP variants, account for 74 percent of variant bases. SNPs dominate by count, while larger changes dominate by DNA real estate. Those larger changes include 292,102 heterozygous insertions and deletions, 559,473 homozygous insertions and deletions reaching up to 82,711 base pairs in length, 53,823 block substitutions, and 90 inversions — the largest stretching 670,000 bases. The team also found 62 copy number variants, regions where whole segments of DNA are gained or lost, 87 percent of which overlapped entries in the Database of Genomic Variants. Additionally, 1,288,319 of the variants were entirely novel — not in dbSNP at the time of analysis. Forty-four percent of protein-coding genes carried at least one heterozygous variant, and 4,107 genes contained nonsynonymous changes — mutations that alter the protein sequence. Between the two chromosomal copies of this one individual, the genome is about 99.5 percent identical. That half-percent difference, it turns out, is far more complex than anyone had appreciated. However, knowing that variants exist at specific positions is different from knowing which variants travel together on the same chromosome. That's haplotype phasing — the ability to say not just that two alleles are present, but which one sits on the maternal copy and which on the paternal. Levy and colleagues tackled this directly using paired-end and mate-pair sequencing data. They started with the 1,856,446 autosomal heterozygous variants and encoded the sequencing information as a matrix, with reads as rows and variant sites as columns. A greedy algorithm seeded haplotype pairs from the most informative rows and iteratively assigned overlapping reads by majority rule. The key ingredient was mate pairing. Single reads couldn't bridge many variant sites — the average spacing between heterozygous variants was about 1,500 base pairs. But mate pairs linked each variant to an average of 8.7 others. The payoff was substantial: haplotypes spanning more than 200 kilobases covered 1.5 gigabases of genome sequence, with 91 percent of variants within those spans correctly assigned. Fewer than one in forty phasings conflicted with independently established HapMap haplotypes in regions of strong linkage disequilibrium. This approach doesn't require family data or population reference panels; it reads the physical chromosome directly. And that matters enormously when you want to connect genotype to phenotype, because it's the combination of variants on a single chromosome that determines what a gene actually does. Which brings us to what this genome says about one specific person. Levy and colleagues ran Venter's diploid sequence against a library of known genotype-phenotype associations, and the results provide a preview of what personalized genomics could actually look like. Some findings were unambiguous. The donor carries 18 and 17 CAG repeats in the Huntington's disease gene — well below the threshold of 29 repeats that defines risk — consistent with no family history and no symptoms. For cardiovascular disease, the picture is more complicated: the donor is heterozygous for two KLOTHO variants previously associated with lower coronary artery disease risk, but simultaneously homozygous at a position in the MMP3 promoter linked to increased risk of acute myocardial infarction. Opposing alleles and competing probabilities is exactly how complex disease works. Several trait associations were also examined. The donor's LCT genotype should, according to published literature, confer adult lactose tolerance. The donor self-reports lactose intolerance — an explicit discordance that Levy and colleagues flag as possibly reflecting other genes or environmental factors. The Clock gene variant rs1801260 is C/C, associated with evening preference in circadian studies. OCA2 variants linked to blue eyes and fair skin are present. The ABCC11 gene variant rs17822931 is G/G, the allele for wet earwax. The donor carries the DRD4 four-repeat allele in exon III, whereas longer forms have been associated with higher novelty seeking. Beyond these trait associations, the assembly also turned up genuinely novel findings — including a four-base-pair heterozygous deletion in the ACOX2 gene predicted to truncate the protein and likely abolish its peroxisomal targeting signal, with unknown but potentially real biological consequences. These examples — the concordances, the discordances, and the open questions — are not flaws in the approach. They represent the honest picture of what one genome can and cannot tell you. A single genotype prediction sits inside a probabilistic cloud of environmental effects, modifier genes, and population frequencies that one genome cannot resolve alone. Levy and colleagues are clear-eyed about the limits. Haplotype phasing remains incomplete in places — phase cannot be determined in as many as 20 percent of cases in this dataset. About 150 megabases of HuRef sequence had no one-to-one alignment to the reference, much of it within segmental duplications. The Y chromosome achieved only about 59 percent coverage. And one genome, however deeply characterized, is not a population sample. What this work establishes is a framework: an allele-aware assembly pipeline, a haplotype construction strategy, and a variant catalog far richer than any SNP-only survey had produced. The comparison to other genomes suggests that genetic variation between two individuals may be up to five times higher than previously estimated. That revision alone reframes how we should think about human diversity. This diploid genome is a starting point — the first molecularly complete portrait of a real person's two-copy genome, and the scaffold on which the era of individualized genomic medicine begins to be built. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

If every human cell contains two copies of every chromosome — one from your mother and one from your father — then any genome sequence that treats you as a single creature with one copy is missing half the picture. If it's missing half the picture, every variant it finds is an undercount, and every variant it misses could be the one that matters for your health. That's the problem Levy and colleagues set out to fix in two thousand seven, when they sequenced and assembled both copies of one person's genome at once for the first time. That person was J. Craig Venter. What they found was more than four million differences from the reference genome, and nearly three-quarters of the variant bases were not the kind anyone had been systematically counting. Here's what had been missing. The Human Genome Project's reference assembly is a composite drawn from multiple donors, and the earlier Celera assembly is a consensus from five individuals. Both approaches collapse two chromosomal copies into one linear sequence, resulting in a mosaic — a patchwork that doesn't exist on any single chromosome in any real person. Single nucleotide polymorphisms, or SNPs, were the main thing these approaches tracked: positions where one letter in the sequence differs. However, a genome is not just a string of single-letter edits. It contains insertions, deletions, duplications, and inversions — large-scale changes that SNP surveys largely ignore.

To see all of that, you need to sequence a real diploid genome and keep the two copies separate. The HuRef project did exactly that. Levy and colleagues generated roughly 32 million sequence reads using Sanger dideoxy technology, with paired-end libraries producing read pairs of about 800 bases from each clone end, achieving approximately 7.5-fold coverage across the genome. The core technical challenge was assembly. Standard assemblers collapse reads from both chromosomal copies into a single consensus, which is precisely what the team wanted to avoid. So they built a modified version of the Celera Assembler that treated each region of variation as a block bounded by nonvariant columns on either side. Reads spanning a variable block were sorted into allele groups; an allele needed at least two supporting reads to be confirmed, and the highest-scoring confirmed allele became the primary consensus while alternate alleles were preserved separately. Think of it as reassembling a shredded book where every page exists in two slightly different editions — the assembler reconstructs both editions rather than averaging them into one that matches neither. The result was 4,528 scaffolds containing 2,810 megabases of contiguous sequence, with a scaffold N50 of 19.5 megabases and roughly 68 percent fewer internal gaps than the earlier whole-genome shotgun assembly.

Now for what that assembly revealed. The total variant catalog came to 4,118,889 filtered events, spanning more than 12 million bases. Of those, 3,213,401 are SNPs — about 78 percent of all events. But here is the finding that changes how you think about human genetic variation: the remaining 22 percent of events, the non-SNP variants, account for 74 percent of variant bases. SNPs dominate by count, while larger changes dominate by DNA real estate. Those larger changes include 292,102 heterozygous insertions and deletions, 559,473 homozygous insertions and deletions reaching up to 82,711 base pairs in length, 53,823 block substitutions, and 90 inversions — the largest stretching 670,000 bases. The team also found 62 copy number variants, regions where whole segments of DNA are gained or lost, 87 percent of which overlapped entries in the Database of Genomic Variants. Additionally, 1,288,319 of the variants were entirely novel — not in dbSNP at the time of analysis. Forty-four percent of protein-coding genes carried at least one heterozygous variant, and 4,107 genes contained nonsynonymous changes — mutations that alter the protein sequence. Between the two chromosomal copies of this one individual, the genome is about 99.5 percent identical. That half-percent difference, it turns out, is far more complex than anyone had appreciated.

However, knowing that variants exist at specific positions is different from knowing which variants travel together on the same chromosome. That's haplotype phasing — the ability to say not just that two alleles are present, but which one sits on the maternal copy and which on the paternal. Levy and colleagues tackled this directly using paired-end and mate-pair sequencing data. They started with the 1,856,446 autosomal heterozygous variants and encoded the sequencing information as a matrix, with reads as rows and variant sites as columns. A greedy algorithm seeded haplotype pairs from the most informative rows and iteratively assigned overlapping reads by majority rule. The key ingredient was mate pairing. Single reads couldn't bridge many variant sites — the average spacing between heterozygous variants was about 1,500 base pairs. But mate pairs linked each variant to an average of 8.7 others. The payoff was substantial: haplotypes spanning more than 200 kilobases covered 1.5 gigabases of genome sequence, with 91 percent of variants within those spans correctly assigned.

Fewer than one in forty phasings conflicted with independently established HapMap haplotypes in regions of strong linkage disequilibrium. This approach doesn't require family data or population reference panels; it reads the physical chromosome directly. And that matters enormously when you want to connect genotype to phenotype, because it's the combination of variants on a single chromosome that determines what a gene actually does. Which brings us to what this genome says about one specific person. Levy and colleagues ran Venter's diploid sequence against a library of known genotype-phenotype associations, and the results provide a preview of what personalized genomics could actually look like. Some findings were unambiguous. The donor carries 18 and 17 CAG repeats in the Huntington's disease gene — well below the threshold of 29 repeats that defines risk — consistent with no family history and no symptoms. For cardiovascular disease, the picture is more complicated: the donor is heterozygous for two KLOTHO variants previously associated with lower coronary artery disease risk, but simultaneously homozygous at a position in the MMP3 promoter linked to increased risk of acute myocardial infarction. Opposing alleles and competing probabilities is exactly how complex disease works.

Several trait associations were also examined. The donor's LCT genotype should, according to published literature, confer adult lactose tolerance. The donor self-reports lactose intolerance — an explicit discordance that Levy and colleagues flag as possibly reflecting other genes or environmental factors. The Clock gene variant rs1801260 is C/C, associated with evening preference in circadian studies. OCA2 variants linked to blue eyes and fair skin are present. The ABCC11 gene variant rs17822931 is G/G, the allele for wet earwax. The donor carries the DRD4 four-repeat allele in exon III, whereas longer forms have been associated with higher novelty seeking. Beyond these trait associations, the assembly also turned up genuinely novel findings — including a four-base-pair heterozygous deletion in the ACOX2 gene predicted to truncate the protein and likely abolish its peroxisomal targeting signal, with unknown but potentially real biological consequences. These examples — the concordances, the discordances, and the open questions — are not flaws in the approach. They represent the honest picture of what one genome can and cannot tell you. A single genotype prediction sits inside a probabilistic cloud of environmental effects, modifier genes, and population frequencies that one genome cannot resolve alone.

Levy and colleagues are clear-eyed about the limits. Haplotype phasing remains incomplete in places — phase cannot be determined in as many as 20 percent of cases in this dataset. About 150 megabases of HuRef sequence had no one-to-one alignment to the reference, much of it within segmental duplications. The Y chromosome achieved only about 59 percent coverage. And one genome, however deeply characterized, is not a population sample. What this work establishes is a framework: an allele-aware assembly pipeline, a haplotype construction strategy, and a variant catalog far richer than any SNP-only survey had produced. The comparison to other genomes suggests that genetic variation between two individuals may be up to five times higher than previously estimated. That revision alone reframes how we should think about human diversity. This diploid genome is a starting point — the first molecularly complete portrait of a real person's two-copy genome, and the scaffold on which the era of individualized genomic medicine begins to be built. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Biochemistry, Genetics and Molecular Biology