Deep Sequencing of the Oral Microbiome Reveals Signatures of Periodontal Disease

Bo Liu, Lina L. Faller, Niels Klitgord, Varun Mazumdar, Mohammad Ghodsi, Daniel D. Sommer, Theodore Gibbons, Todd J. Treangen, Yi-Chien Chang, Shan Li, O. Colin Stine, Hatice Hastürk, Simon Kasif, Daniel Segrè, Mihai Pop, Salomon AmarView original
OverviewBalancedalloy voice
If you've ever been told that gum disease is "just in your mouth," here's the twist: periodontitis is woven into the rest of your body. It travels with cardiovascular risk, with diabetes, and with inflammation you feel elsewhere. So, the question isn't just which microbe shows up when your gums bleed. It's what happens to the whole ecosystem under your gumline as it tips from health to disease. Liu and colleagues went after that bigger picture. They didn't chase a single villain. They sequenced the entire community living in those periodontal pockets. They worked with a small but carefully controlled set of 15 subgingival plaque samples—eight diseased sites from two patients with periodontitis and seven healthy sites from three controls. They did it two ways at once. One pass targeted the 16S ribosomal gene to map who is there; the other was shotgun metagenomics, which reads everything it can to capture what those microbes can do. The 16S run on the 454 platform produced about four hundred ninety-five thousand sequences—roughly thirty thousand reads per sample after quality control. The shotgun run on Illumina GAII was massive—approximately two hundred seventy-three million paired seventy-six base reads with libraries around two hundred bases—but here's the catch. Most of that DNA was human. On average, eighty-seven percent of the reads were host, leaving about twelve point four percent—roughly thirty-four million reads—to tell the microbial story. Clearing out the human signal was step one. They mapped reads to the human reference genome with Bowtie, allowing a few mismatches so they wouldn't miss near matches; if either end of a pair hit human, they tossed the pair. Then, they chased down stragglers by comparing the leftovers to a protein database and labeling anything that looked too human with a strict identity cutoff. Only when the host haystack was swept away did they start looking for microbial needles. On the 16S side, they used a conservative approach—cluster sequences into operational taxonomic units, assign names only when a read is a near-exact match, and leave the unknowns unforced. On the shotgun side, they mapped reads to enzyme and pathway annotations so they could ask a different question: even if the names differ, do diseased communities run the same metabolic playbook? Taxonomically, the contrast was crisp. Healthy sites leaned heavily gram-positive—think Streptococcus, Actinomyces, Granulicatella—while disease shifted the center of gravity to gram-negative genera like Prevotella, Treponema, Leptotrichia, and Fusobacterium. That gram-negative tilt wasn't subtle; a simple contingency test across the community made it vanishingly unlikely to be random, with a p-value on the order of ten to the negative fifteenth. One group in particular, the enigmatic TM7 division, stood out. TM7 showed up in eleven of the fifteen samples at more than two percent abundance, averaging about five point seven percent by the 16S survey. Shotgun estimates were in the same ballpark, and in one diseased site, TM7 climbed as high as roughly twenty-seven percent. Across the board, TM7 was enriched in disease, a pattern that cleared the usual significance bar. And there was a canary in the coal mine: one "healthy" tooth, H31, didn't cluster with the other healthy sites at all. It sat with the diseased group in the analyses. A clinical recheck found incipient signs of disease at that site. Now step back from names and look at the shape of the community. The diseased samples weren't just different from healthy; they were similar to each other. In several independent views—principal components based on 16S composition, on shotgun-derived enzyme counts, and even on short DNA word frequencies—the disease points pulled together into a tight island. Healthy points? They spread out, person to person. There's a paradox hiding here that I think is fascinating. Diversity by one measure—Shannon diversity of 16S reads—was actually higher in disease. More kinds of bacteria. Yet, the overall configuration of a diseased community, in both taxonomic and functional space, was more constrained. It's as if disease acts like a magnet in microbiome space, pulling different starting communities into a narrow attractor. The functional readout tells you why that attractor might exist. Shotgun reads mapped to enzymes and pathways told a consistent story: diseased sites are optimized for life in an inflamed, oxygen-poor niche, where they can feast on host-derived nutrients and push back against stress. Pathways for breaking down fatty acids and acetyl-CoA were elevated. So were routes for degrading aromatic amino acids, and redox systems that thrive anaerobically, like ferredoxin-linked processes. Transporters of the energy-coupling factor family—little molecular hoists that pull vitamins and metals into the cell—were also more common. Layer on top of that a suite of functions you'd expect in a community that needles the host. Lipopolysaccharide, the business end of gram-negative endotoxin, was more prominent. So were mobile elements like conjugative transposons and type four secretion components, the machinery microbes use to move genes and talk to eukaryotic cells. Resistance features popped too—especially to metals like mercury, cobalt, zinc, and cadmium. Antibiotic resistance genes were present in both states; metal resistance skewed toward disease. Healthy sites didn't look like a mirror image. They had their own preferences: more of the machinery to build fatty acids than to burn them, along with purine metabolism and glycerol-3-phosphate pathways. Homoserine metabolism, tied into bacterial quorum sensing, was enriched in health; the flip side of that coin is a reduction in quorum-signaling routes in disease, a hint that the social rules of the community change as the pocket deepens. These aren't one-off blips. When Liu's team ranked enzymes and ran Gene Set Enrichment Analysis, or GSEA, across hundreds of functional categories, the disease group clustered tightly in enzyme space, and a consistent suite of pathways rose to the top. They anchored significance with thresholds that protect against false discovery, but the headline isn't the p-values. It's the coherence. Separate people. Shared metabolism. How did they compare such different data types on even footing? For the functional view, they mapped shotgun reads to Kyoto Encyclopedia of Genes and Genomes, or KEGG, Orthology codes, rolled those up to enzyme identities, and used principal components to visualize how samples relate. For broader categorical shifts, they used a pathway enrichment framework—GSEA—configured to shuffle gene sets rather than phenotypes, which is better when sample sizes are small. They also did a clever sanity check that ignores names entirely: a tetramer analysis. Count every four-letter DNA word in a sample, estimate its expected frequency with a simple Markov model—roughly, the probability of a four-letter word is approximated from the two adjoining three-letter words, adjusted by the shared overlap—and then ask how each sample deviates from expectation. Disease samples from different people aligned along the diagonal in that space, a signature of similarity. Healthy samples wandered. All of this would be less convincing if they couldn't connect "who" to "what." That's where the assembly work matters. With host DNA drowning out most microbial reads and no single assembler built for metagenomes at the time, they took a hybrid route. Start with a de novo assembly to stitch together everything you can. Overlay those contigs on a curated oral reference collection to pull in related fragments. Merge, scaffold, and iterate. Across the dataset, that strategy pushed contiguity way up—on average about four times higher N50 than de novo alone, and roughly double what pure reference-guided assembly achieved—with contigs reaching into the kilobase range even at modest depth. The poster child was TM7a, a lineage that resists cultivation. By combining single-cell genome fragments for TM7a with metagenomic contigs and scaffolding with mate-pair links, they expanded the TM7a assembly from about one point seven megabases to roughly two point three, and they doubled the typical contig size from a few hundred bases to nearly eight hundred. That expansion wasn't just more of the same. They recovered seven hundred three genes that hadn't been seen in the single-cell dataset, including a drug-resistance transporter of the EmrB and QacA family, a couple of phage proteins, and housekeeping genes that fill in basic functions. On the flip side of disease, health had its own genomic signature in a familiar genus. Actinomyces, a stalwart of healthy plaque, showed person-specific genomes that looked like cousins, not clones, of the lab reference strain MG1. In two healthy individuals, contigs aligned to MG1 at around ninety-six and ninety-five percent average identity, but with telling differences. Some regions were simply gone—deletions that included a mercury-resistance locus and other sequences linked to virulence in related strains. Other regions were noisier than average, with a high density of single-nucleotide variants in genes that bacteria use to sense and adapt to their environment, like transcriptional regulators and ATP-binding cassette, or ABC, transporters. One especially polymorphic stretch sat in the GAPDH protein, a surface-exposed protein that helps these bacteria stick to host tissues and has been implicated in periodontal colonization. The portrait that emerges is a healthy state that tolerates, maybe even cultivates, individual variation among its gram-positive residents. They didn't eyeball those variants. They mapped reads back to assembled contigs, called changes only where there was support from multiple reads at a depth of at least five, and estimated genome-wide diversity with a model that corrects for the known sequencing error rate. In plain terms, they grouped sites by how deeply they were covered, subtracted out the number of apparent variants you'd expect from errors alone—using an error rate of about one percent as a baseline—and then summarized the true diversity with a sliding window one kilobase wide, stepping a hundred bases at a time. Regions that rose more than two standard deviations above the background were flagged as hotspots. It's careful, conservative, and importantly, it agrees with the broader story: healthy Actinomyces genomes are similar but idiosyncratic; disease genomes converge not only in who's there but in what they carry. Now, this is a pilot study involving five people—a cross-section in time. That heavy host signal is a real limitation; when nearly nine out of ten shotgun reads are human, you are working with a thin microbial slice. Even so, the stack of evidence points the same way. Taxonomy pivots from gram-positive to gram-negative in disease, with TM7 expansion that can reach a quarter of the community in a site. Function pivots toward pathways that burn host-derived lipids and amino acids, exchange genes, and resist metals. Community structure tightens into a shared metabolic neighborhood, even as the list of species lengthens. Health, in contrast, looks more like a set of personalized equilibria—still patterned, but more varied. Why should you care about that shape? Because it's not just descriptive; it's actionable. The one healthy-looking tooth that fell into the disease cluster, H31, foreshadows a clinical transition before full symptoms arrive. That's the diagnostic promise of a systems-ecology view. Instead of chasing "usual suspects" microbe by microbe, you read the configuration: the mix of taxa, the metabolic tuning, the mobile elements that suggest a community ready to exploit a disrupted host. Liu and colleagues showed that you can get there with a dual sequencing strategy, rigorous host filtering, and a hybrid assembly pipeline, even when microbial reads are a minority. Where does this go next? Longitudinal tracking is the big one—watching mouths move through health, dysbiosis, and back again while layering in host inflammation markers to see how the two systems dance. Deeper sequencing will sharpen assemblies and bring rare functions into focus. But even now, the main arc is clear. Periodontal disease isn't a single pathogen stepping into a vacuum. It's a state change, a constrained attractor where different communities settle into a common, parasitic metabolism aligned with inflammation. Once you see it that way, you stop asking "which bug did it?" and start asking "which way is the system leaning, and can we nudge it back?"

If you've ever been told that gum disease is "just in your mouth," here's the twist: periodontitis is woven into the rest of your body. It travels with cardiovascular risk, with diabetes, and with inflammation you feel elsewhere. So, the question isn't just which microbe shows up when your gums bleed.

It's what happens to the whole ecosystem under your gumline as it tips from health to disease. Liu and colleagues went after that bigger picture. They didn't chase a single villain. They sequenced the entire community living in those periodontal pockets.

They worked with a small but carefully controlled set of 15 subgingival plaque samples—eight diseased sites from two patients with periodontitis and seven healthy sites from three controls. They did it two ways at once. One pass targeted the 16S ribosomal gene to map who is there; the other was shotgun metagenomics, which reads everything it can to capture what those microbes can do.

The 16S run on the 454 platform produced about four hundred ninety-five thousand sequences—roughly thirty thousand reads per sample after quality control. The shotgun run on Illumina GAII was massive—approximately two hundred seventy-three million paired seventy-six base reads with libraries around two hundred bases—but here's the catch. Most of that DNA was human.

On average, eighty-seven percent of the reads were host, leaving about twelve point four percent—roughly thirty-four million reads—to tell the microbial story.

Clearing out the human signal was step one. They mapped reads to the human reference genome with Bowtie, allowing a few mismatches so they wouldn't miss near matches; if either end of a pair hit human, they tossed the pair. Then, they chased down stragglers by comparing the leftovers to a protein database and labeling anything that looked too human with a strict identity cutoff.

Only when the host haystack was swept away did they start looking for microbial needles. On the 16S side, they used a conservative approach—cluster sequences into operational taxonomic units, assign names only when a read is a near-exact match, and leave the unknowns unforced. On the shotgun side, they mapped reads to enzyme and pathway annotations so they could ask a different question: even if the names differ, do diseased communities run the same metabolic playbook?

Taxonomically, the contrast was crisp. Healthy sites leaned heavily gram-positive—think Streptococcus, Actinomyces, Granulicatella—while disease shifted the center of gravity to gram-negative genera like Prevotella, Treponema, Leptotrichia, and Fusobacterium. That gram-negative tilt wasn't subtle; a simple contingency test across the community made it vanishingly unlikely to be random, with a p-value on the order of ten to the negative fifteenth.

One group in particular, the enigmatic TM7 division, stood out. TM7 showed up in eleven of the fifteen samples at more than two percent abundance, averaging about five point seven percent by the 16S survey. Shotgun estimates were in the same ballpark, and in one diseased site, TM7 climbed as high as roughly twenty-seven percent.

Across the board, TM7 was enriched in disease, a pattern that cleared the usual significance bar. And there was a canary in the coal mine: one "healthy" tooth, H31, didn't cluster with the other healthy sites at all. It sat with the diseased group in the analyses. A clinical recheck found incipient signs of disease at that site.

Now step back from names and look at the shape of the community. The diseased samples weren't just different from healthy; they were similar to each other. In several independent views—principal components based on 16S composition, on shotgun-derived enzyme counts, and even on short DNA word frequencies—the disease points pulled together into a tight island.

Healthy points? They spread out, person to person. There's a paradox hiding here that I think is fascinating.

Diversity by one measure—Shannon diversity of 16S reads—was actually higher in disease. More kinds of bacteria. Yet, the overall configuration of a diseased community, in both taxonomic and functional space, was more constrained.

It's as if disease acts like a magnet in microbiome space, pulling different starting communities into a narrow attractor.

The functional readout tells you why that attractor might exist. Shotgun reads mapped to enzymes and pathways told a consistent story: diseased sites are optimized for life in an inflamed, oxygen-poor niche, where they can feast on host-derived nutrients and push back against stress. Pathways for breaking down fatty acids and acetyl-CoA were elevated.

So were routes for degrading aromatic amino acids, and redox systems that thrive anaerobically, like ferredoxin-linked processes. Transporters of the energy-coupling factor family—little molecular hoists that pull vitamins and metals into the cell—were also more common. Layer on top of that a suite of functions you'd expect in a community that needles the host.

Lipopolysaccharide, the business end of gram-negative endotoxin, was more prominent. So were mobile elements like conjugative transposons and type four secretion components, the machinery microbes use to move genes and talk to eukaryotic cells. Resistance features popped too—especially to metals like mercury, cobalt, zinc, and cadmium.

Antibiotic resistance genes were present in both states; metal resistance skewed toward disease.

Healthy sites didn't look like a mirror image. They had their own preferences: more of the machinery to build fatty acids than to burn them, along with purine metabolism and glycerol-3-phosphate pathways. Homoserine metabolism, tied into bacterial quorum sensing, was enriched in health; the flip side of that coin is a reduction in quorum-signaling routes in disease, a hint that the social rules of the community change as the pocket deepens.

These aren't one-off blips. When Liu's team ranked enzymes and ran Gene Set Enrichment Analysis, or GSEA, across hundreds of functional categories, the disease group clustered tightly in enzyme space, and a consistent suite of pathways rose to the top. They anchored significance with thresholds that protect against false discovery, but the headline isn't the p-values. It's the coherence. Separate people. Shared metabolism.

How did they compare such different data types on even footing? For the functional view, they mapped shotgun reads to Kyoto Encyclopedia of Genes and Genomes, or KEGG, Orthology codes, rolled those up to enzyme identities, and used principal components to visualize how samples relate. For broader categorical shifts, they used a pathway enrichment framework—GSEA—configured to shuffle gene sets rather than phenotypes, which is better when sample sizes are small.

They also did a clever sanity check that ignores names entirely: a tetramer analysis. Count every four-letter DNA word in a sample, estimate its expected frequency with a simple Markov model—roughly, the probability of a four-letter word is approximated from the two adjoining three-letter words, adjusted by the shared overlap—and then ask how each sample deviates from expectation. Disease samples from different people aligned along the diagonal in that space, a signature of similarity. Healthy samples wandered.

All of this would be less convincing if they couldn't connect "who" to "what." That's where the assembly work matters. With host DNA drowning out most microbial reads and no single assembler built for metagenomes at the time, they took a hybrid route. Start with a de novo assembly to stitch together everything you can.

Overlay those contigs on a curated oral reference collection to pull in related fragments. Merge, scaffold, and iterate. Across the dataset, that strategy pushed contiguity way up—on average about four times higher N50 than de novo alone, and roughly double what pure reference-guided assembly achieved—with contigs reaching into the kilobase range even at modest depth.

The poster child was TM7a, a lineage that resists cultivation. By combining single-cell genome fragments for TM7a with metagenomic contigs and scaffolding with mate-pair links, they expanded the TM7a assembly from about one point seven megabases to roughly two point three, and they doubled the typical contig size from a few hundred bases to nearly eight hundred. That expansion wasn't just more of the same.

They recovered seven hundred three genes that hadn't been seen in the single-cell dataset, including a drug-resistance transporter of the EmrB and QacA family, a couple of phage proteins, and housekeeping genes that fill in basic functions.

On the flip side of disease, health had its own genomic signature in a familiar genus. Actinomyces, a stalwart of healthy plaque, showed person-specific genomes that looked like cousins, not clones, of the lab reference strain MG1. In two healthy individuals, contigs aligned to MG1 at around ninety-six and ninety-five percent average identity, but with telling differences.

Some regions were simply gone—deletions that included a mercury-resistance locus and other sequences linked to virulence in related strains. Other regions were noisier than average, with a high density of single-nucleotide variants in genes that bacteria use to sense and adapt to their environment, like transcriptional regulators and ATP-binding cassette, or ABC, transporters. One especially polymorphic stretch sat in the GAPDH protein, a surface-exposed protein that helps these bacteria stick to host tissues and has been implicated in periodontal colonization.

The portrait that emerges is a healthy state that tolerates, maybe even cultivates, individual variation among its gram-positive residents.

They didn't eyeball those variants. They mapped reads back to assembled contigs, called changes only where there was support from multiple reads at a depth of at least five, and estimated genome-wide diversity with a model that corrects for the known sequencing error rate. In plain terms, they grouped sites by how deeply they were covered, subtracted out the number of apparent variants you'd expect from errors alone—using an error rate of about one percent as a baseline—and then summarized the true diversity with a sliding window one kilobase wide, stepping a hundred bases at a time.

Regions that rose more than two standard deviations above the background were flagged as hotspots. It's careful, conservative, and importantly, it agrees with the broader story: healthy Actinomyces genomes are similar but idiosyncratic; disease genomes converge not only in who's there but in what they carry.

Now, this is a pilot study involving five people—a cross-section in time. That heavy host signal is a real limitation; when nearly nine out of ten shotgun reads are human, you are working with a thin microbial slice. Even so, the stack of evidence points the same way.

Taxonomy pivots from gram-positive to gram-negative in disease, with TM7 expansion that can reach a quarter of the community in a site. Function pivots toward pathways that burn host-derived lipids and amino acids, exchange genes, and resist metals. Community structure tightens into a shared metabolic neighborhood, even as the list of species lengthens.

Health, in contrast, looks more like a set of personalized equilibria—still patterned, but more varied.

Why should you care about that shape? Because it's not just descriptive; it's actionable. The one healthy-looking tooth that fell into the disease cluster, H31, foreshadows a clinical transition before full symptoms arrive.

That's the diagnostic promise of a systems-ecology view. Instead of chasing "usual suspects" microbe by microbe, you read the configuration: the mix of taxa, the metabolic tuning, the mobile elements that suggest a community ready to exploit a disrupted host. Liu and colleagues showed that you can get there with a dual sequencing strategy, rigorous host filtering, and a hybrid assembly pipeline, even when microbial reads are a minority.

Where does this go next? Longitudinal tracking is the big one—watching mouths move through health, dysbiosis, and back again while layering in host inflammation markers to see how the two systems dance. Deeper sequencing will sharpen assemblies and bring rare functions into focus.

But even now, the main arc is clear. Periodontal disease isn't a single pathogen stepping into a vacuum. It's a state change, a constrained attractor where different communities settle into a common, parasitic metabolism aligned with inflammation.

Once you see it that way, you stop asking "which bug did it?" and start asking "which way is the system leaning, and can we nudge it back?"

More in Dentistry