Defining the oral microbiome by whole-genome sequencing and resistome analysisthe complexity of the healthy picture
Think about your mouth as a small, bustling city with hundreds of neighborhoods and thousands of residents. Some are permanent, while some just pass through with your morning coffee. Scientists have counted more than seven hundred microbial species that can live there, including bacteria, fungi, viruses, archaea, and even protozoa.
And yet, even with all that attention, a third of this city remains uncultivated, and only about half of its residents have names and faces we recognize from culture. Most of what we know comes from bacterial surveys, usually by reading a marker gene called sixteen S ribosomal RNA. That gave us the skyline, but it missed the side streets — strains within species and anything that isn’t a bacterium.
Caselli and colleagues set out to change that picture. Their idea was simple and ambitious: don’t just sample the city; map it block by block across kingdoms and sketch the city’s arsenal — its antibiotic resistance genes — while you’re at it.
So here’s what they did. They recruited twenty healthy young adults in Northern Italy — ten men and ten women with an average age of around twenty-five — and sampled eight micro habitats in each mouth. They did not just collect saliva but also an oral rinse after swishing with buffered saline, the tongue dorsum, the hard palate, the buccal mucosa, the keratinized gingiva, and two kinds of dental plaque: supragingival and subgingival.
Then they read everything, not just a marker gene. They conducted whole-genome sequencing on an Ion Torrent S5, using Kraken2 to assign reads against a database covering archaea, bacteria, fungi, protozoa, and viruses. They set a conservative threshold — at least ten copies were required to call a genus or species — and proceeded to count who was there.
In parallel, they profiled the resistome — the collection of antibiotic resistance genes — using a real-time quantitative polymerase chain reaction microarray that targets eighty-four genes spanning major drug classes. Power calculations suggested that twenty people were enough to reliably detect differences among key genera. The team is clear about the caveats: young adults, one geography, and a resistome assay that can’t tie a specific gene to a specific microbe in this design.
The big picture that pops out is this: every habitat in the mouth has its own community, and those habitat signatures are stronger than person-to-person differences. When the team measured alpha diversity — think of it as “how many kinds and how evenly spread” — subjects looked broadly similar as a group, with Shannon index values between about 3.2 and 5.4, and there was no significant difference across individuals. By contrast, sites varied dramatically.
Subgingival plaque sat at the high end with an index around 7.0. Keratinized gingiva was at the low end, roughly 1.9. Saliva, tongue, and supragingival plaque fell in the middle, forming six distinct clusters overall, and those site effects weren’t subtle — the statistics pegged them as highly significant.
There was a small wrinkle: people who used manual toothbrushes showed a bit more variability in diversity than powered brush users, with a p-value around 0.02. This is interesting, but the habitat story still dominates.
Diversity is one lens, and composition is another. When they looked at beta diversity — how different whole communities are from each other — the site stamps were unmistakable. Weighted UniFrac, a metric that considers both presence and relatedness of taxa, separated the eight micro habitats cleanly.
Some pairs made intuitive sense: supragingival and subgingival plaque were cousins; oral mucosa and keratinized gingiva clustered together; saliva and oral rinse marched in the same direction. But saliva and rinse also tracked, to varying degrees, with the underlying site-specific samples, a clue we’ll come back to when we talk about practical sampling.
To put numbers on which microbes define each place, Caselli and colleagues trained a site classifier using a technique called Prediction Analysis for Microarrays, or PAM for short. It’s essentially a nearest shrunken centroid method that highlights the most informative features. Sixty-two genera emerged as the best markers of place.
Feed a sample into the model, and it called the site correctly three-quarters of the time. Some habitats were easy to spot — tongue, oral mucosa, supragingival plaque, and keratinized gingiva hit eighty to ninety percent accuracy. Others blurred at the edges: hard palate and subgingival plaque hovered in the fifty-five to sixty percent range.
That distribution tells you something about how distinct those communities are and where the boundaries in this city are sharp versus fuzzy.
Who are the dominant residents? On mucosal surfaces, Streptococci are the major landlords. The hard palate came in with about forty-four percent Streptococci, oral mucosa roughly sixty-five percent, and keratinized gingiva around sixty-six percent.
Move over to tongue, saliva, and plaque, and that Streptococcal share dropped into the low twenties or teens. Zooming to species, Streptococcus mitis stood out as the single most prevalent, at around nine-point-five percent overall, with S. oralis, S. salivarius, and S. sanguinis close behind. Among the non-Streptococcal cast, Haemophilus parainfluenzae made a strong showing at roughly eleven-point-eight percent.
And while the study focused on a cross-kingdom view, the non-bacterial members were present in small but detectable amounts: fungi accounted for about zero-point-zero-zero-four percent of counts, largely Candida species, more noticeable on the hard palate and in supragingival plaque and oral rinse. Viruses tallied roughly zero-point-zero-three percent of counts, mainly bacteriophages — families like Siphoviridae, Myoviridae, and Podoviridae — with a trace of human Herpesviruses; Herpesvirales sat around six ten-thousandths of the total.
That’s not much by mass, but it matters ecologically. Phages shape bacterial populations, and even a small human virome can have outsize effects.
Now, if you could only collect one sample from someone — and that’s often the case in big studies — which should it be? Saliva is easy. But this study points to a better proxy.
Oral rinse, the sample you get after a one-minute swish with buffered saline, captured the whole-mouth picture more faithfully than unstimulated saliva. When Caselli and colleagues compared the average composition of oral rinse to the combined profile of all site-specific samples, the match was striking, with an analysis of variance similarity near zero-point-nine-eight. Saliva lagged, closer to zero-point-seven.
You could see the same pattern in the taxa driving the PAM classifier and in side-by-side abundance comparisons: rinse mirrored the mosaic; saliva emphasized a narrower slice. If your goal is surveillance or establishing baselines across populations, that matters. One minute of rinse buys you a much more representative snapshot of the city.
Let’s turn to the city’s arsenal. Even healthy mouths carry resistance genes, and this study provides a clean baseline for what “normal” looks like in young adults. Using a quantitative polymerase chain reaction microarray on the oral rinse DNA, the team looked for eighty-four resistance markers across major drug classes — macrolide-lincosamide-streptogramin, often abbreviated as MLS, tetracyclines, beta-lactams, fluoroquinolones, glycopeptides like vancomycin, and more.
Two classes dominated: MLS and tetracyclines. If you compare to negative controls and take the logarithm base ten of the fold change, three genes stood well above the rest: the mefA gene at about five-point-four-three, the ermB gene around four-point-six, and the tetB gene near four-point-one-nine. Put plainly, those determinants were abundant.
Others — the ermA gene, the SHV-group beta-lactamases, the VIM-1 metallo-beta-lactamases, quinolone resistance genes in the Qnr family, the msrA efflux gene, the mecA gene associated with methicillin resistance, and the vanC gene — were present at lower levels. There were gender differences for several of these rarer genes, with p-values under zero-point-zero-one, and a pattern where some markers like the aac gene, SHV, VIM-1, Qnr, msrA, and mecA tended to be higher in males, while the aac2 gene and vanC leaned higher in females. It’s an intriguing signal, one that needs larger cohorts to interpret, but it reminds us that a shared baseline can still carry individualized features.
A quick word on what those resistance numbers mean. The microarray measures gene presence and relative abundance in the total DNA from the rinse — one microgram per reaction in this study — but it doesn’t assign a particular gene to a particular species. So we learn that the healthy oral ecosystem has the potential for resistance spread, and which mechanisms are common, but not which microbe is carrying what.
That’s a design choice here, made to get a wide scan across drug classes. Whole-genome sequencing could, in principle, resolve gene-host links but would demand deeper coverage and more complex assembly than was practical across all samples.
Methodologically, a couple of details are worth calling out because they influence what you can see. Before extracting DNA, the team used a lysozyme pre-lysis step to crack open Gram-positive bacteria more effectively. That helps capture tough-walled residents like many Streptococci and Actinomyces.
For the sequencing runs, they standardized input to one hundred nanograms of DNA per sample and set a ten-copy detection threshold for calling a taxon present, a move that reduces false positives from stray reads. Between those guardrails and the cross-kingdom reference database, they were able to detect at least two hundred eighteen genera and five hundred seventy species across the cohort. That’s a lot of city blocks and enough resolution to see both a common core and local zoning laws.
It’s also not a free lunch. Whole metagenome work has trade-offs. It’s richer than sixteen S because you see strains and non-bacterial players, and you can infer functions.
But it’s hungry — for sequencing depth, for compute, and for careful curation of reference databases. Caselli and colleagues were transparent about those limits. Add in the demographics — all twenty participants were young adults from the same region — and you have to be careful about generalizing to older populations, different diets, and different geographies.
That said, their power analysis suggested a zero-point-nine-six chance of detecting targeted genus-level differences with twenty subjects, and the internal consistency of the habitat effects backs that up. The key point is the baseline: a coherent, site-specific map of what health looks like.
What do we do with that map? Two practical takeaways stand out. First, if you’re designing large studies or public health surveillance and you can only take one sample, oral rinse gives you a more representative cross-section than saliva.
It’s quick, it’s scalable, and in this dataset, it aligned closely with the aggregate of all eight sites. Second, when you see resistance genes in a clinical sample, context matters. Healthy mouths carry mefA, ermB, and tetB at appreciable levels.
That doesn’t mean a person has a resistant infection; it means the oral ecosystem contains those capacities, likely distributed across commensals. Having a quantitative baseline lets clinicians and researchers spot when a resistome is truly shifted — and in which direction.
Stepping back, there’s a satisfying symmetry here. A decade ago, sixteen S told us who the big players were and that dental plaque wasn’t the same as tongue. Now, whole-genome sequencing fills in the faces and the fine print across bacteria, fungi, and viruses, and ties those portraits to the habitats they prefer.
Caselli and colleagues showed that in twenty mouths, taken one neighborhood at a time, a city emerges: distinct districts, familiar landmarks, and a shared civic character. It’s a living map. And like any good map, it’s most powerful when you use it — to guide sampling choices, to interpret perturbations, and to decide when a detour from health is a pothole or a sinkhole.
The next chapters — older cohorts, different regions, linking resistance genes to hosts — will add street names and traffic patterns. But the grid is there now, clear and navigable, and that’s a leap.
Related lectures
- Comparison of oral microbiota in tumor and non-tumor tissues of patients with oral squamous cell carcinoma
- Beyond Streptococcus mutans: Dental Caries Onset Linked to Multiple Species by 16S rRNA Community Analysis
- Oral pathobiont induces systemic inflammation and metabolic changes associated with alteration of gut microbiota
- Streptococcus mutans-derived extracellular matrix in cariogenic oral biofilms
- Oral Biofilm Architecture on Natural Teeth
- The salivary microbiota as a diagnostic indicator of oral cancer: A descriptive, non-randomized study of cancer-free and oral squamous cell carcinoma subjects