Intestinal dysbiosis in preterm infants preceding necrotizing enterocolitisa systematic review and meta-analysis

Mohan Pammi, Julia Cope, Phillip I. Tarr, Barbara Warner, Ardythe L. Morrow, Volker Mai, Katherine E. Gregory, J. Simon Kroll, Valerie McMurtry, Michael J. Ferris, Lars Engstrand, Heléne Engstrand Lilja, Emily B. Hollister, James Versalovic, Josef NeuView original
OverviewBalancedalloy voice
If you've ever set foot in a neonatal intensive care unit, you know it’s a place where hours stretch and every gram matters. In that world, necrotizing enterocolitis, or NEC, is the diagnosis everyone dreads. Among very low birthweight babies, those under 1500 grams, about seven percent will face it, and it accounts for up to five percent of neonatal intensive care admissions. Mortality sits in a brutal band of approximately fifteen to thirty percent, and survivors can carry neurodevelopmental scars for years. That’s why a simple idea has had such staying power: maybe before NEC erupts, the gut ecosystem tips out of balance. Dysbiosis. Fewer friendly neighbors, more potentially harmful ones. Animal studies hint at this—germ-free pups don’t get NEC—and human studies point in the same direction, but they’ve been small, scattered, and sometimes contradictory. Pammi and colleagues decided to stop reading the tea leaves one cup at a time and brew a pot. Their question was sharp: can we define a microbial pattern that reliably precedes NEC in preterm infants, and can we explain why past studies disagree? They built a harmonized evidence base with tight study selection, uniform reprocessing of raw sequence data, and a plan to account for the usual suspects that muddy microbiome studies: antibiotics, feeding, delivery mode, and even which region of the 16S ribosomal RNA gene was sequenced. The goal wasn’t just to pool data. It was to compare apples with apples. The search was broad and methodical. Using the Cochrane Neonatal Review Group strategy, they sifted through six thousand eight hundred and twelve records, screened four thousand two hundred and seventeen after de-duplication, and pulled twenty-four for full-text review. Fourteen met inclusion criteria. Of those, nine teams shared not just summary results but the raw 16S reads plus clinical metadata. One study had to be set aside because the sequencing was too shallow. That left eight studies in the quantitative synthesis, involving one hundred and six infants who developed NEC, two hundred and seventy-eight controls, and a total of two thousand nine hundred and forty-four stool samples. If you’re picturing a lot of tiny diapers and time points, you’re right—that’s the strength here. Many repeated samples per infant allow you to watch the microbiome evolve toward, or away from, disease. To keep the analysis clean, every sequence ran through the same pipeline. They used Quantitative Insights Into Microbial Ecology, or QIIME, version one point eight point zero, which is a standard tool for microbial community analysis, and assigned operational taxonomic units, or OTUs, basically species bins, by matching reads to the GreenGenes reference. Only high-quality reads were kept: average quality scores had to be at least twenty-five, reads needed to be between two hundred and one thousand bases, barcodes and primers had to match exactly, and there couldn’t be ambiguous bases or long homopolymers. Samples with fewer than one thousand two hundred reads were dropped, which is why that one study didn’t make the cut. Then they rarefied, essentially down-sampling so every sample had the same number of reads. That levels the playing field. On timing, the diagnosis of NEC landed where clinicians would expect: a mean corrected gestational age—age adjusted for prematurity—of thirty point one weeks with a spread of about two point four weeks in the sixty-one infants where that date was available. Here’s the headline. Before NEC shows itself clinically, the gut tilts in a consistent direction. Across corrected gestational ages, infants who went on to develop NEC had more Proteobacteria and fewer Firmicutes and Bacteroidetes compared with controls. Think of Proteobacteria as a big tent that includes families like Enterobacteriaceae—Klebsiella and its cousins—and Firmicutes and Bacteroidetes as many of the usual early-life commensals. That tilt wasn’t random jitter; it began earlier and seemed to gather pace around twenty-seven weeks corrected age, approaching the thirty-week window when NEC typically hits. When Pammi and colleagues tested those phylum-level differences across all corrected ages, Proteobacteria rose and Firmicutes and Bacteroidetes fell with p-values below zero point zero five. The big picture is a directional shift, not a single smoking-gun microbe. Now, you might ask: did overall diversity collapse before NEC? Surprisingly, not in a way you could rely on. When they measured alpha diversity—how many kinds of bacteria are present and how evenly they’re represented—using observed OTUs, the Shannon index, and the Simpson index, there weren’t consistent case–control differences across the data. In a model tailored to count data, species richness did climb with gestational maturity—older guts harbor more kinds of microbes—with very strong support. After controlling for age, NEC cases trended toward lower richness than controls, but that signal just missed conventional significance. And when they looked at beta diversity—how different one sample is from another—using UniFrac, a distance that accounts for how related the bacteria are on the tree of life, the principal coordinates plots didn’t split into neat NEC versus control clouds. Composition changed, but generic diversity metrics alone didn’t draw a diagnostic line. Why so much fuzz? Methodology, for one. Not all 16S amplicons are created equal. Three of the included studies sequenced the V1 to V3 hypervariable region of the 16S gene, while five targeted V3 to V5. Alpha diversity didn’t change much based on region—that is, counting species didn’t depend on which region you sequenced—but the relative proportions of phyla did. The V3 to V5 studies tended to report more Proteobacteria and fewer Firmicutes than V1 to V3 studies. You can think of it as primer and region choice turning the lens slightly, making some groups look bigger and others smaller. In fact, even looking within controls, Shannon diversity differed when you compared V1 to V3 to V3 to V5, and that held when comparing NEC samples sequenced with V3 to V5 to controls sequenced with V1 to V3. Study-level quirks cropped up too. Work by Normann, which used V3 to V5, and by Torraza, which used V1 to V3, leaned toward higher Firmicutes and lower Proteobacteria than the rest, reminding us that no pipeline erases all variability. Clinical exposures also pulled hard on the microbiome. Antibiotics, unsurprisingly, left a deep footprint. In this meta-analysis, there were roughly one thousand eight hundred control samples collected during antibiotic exposure and just over one hundred collected without it; among NEC samples, about two hundred seventy were during antibiotics and a few dozen without. Diversity shifted with antibiotics—OTU richness and the Shannon index differed significantly between exposed and unexposed controls—and the overall community structure clustered by exposure when you looked at UniFrac distances. Taxonomically, antibiotics pushed the gut toward Proteobacteria and away from Firmicutes, Actinobacteria, and Bacteroidetes. Zooming in to the genus level among controls, exposure was linked to higher abundances of Klebsiella, unclassified Enterobacteriaceae, Proteus, Paenibacillus, Epulopiscium, and Pseudomonas. In antibiotic-free windows, Clostridium and unclassified Clostridiaceae were relatively more common. None of that is shocking if you’ve seen what broad-spectrum drugs do to a microbial community. It matters here because it can masquerade as, or amplify, the pre-NEC signature. Feeding and delivery mode added nuance rather than clean separation. Looking within controls, formula-fed infants tended to have more Firmicutes, while those on breast milk showed higher relative Proteobacteria. It’s a reminder that human milk comes with its own microbes and oligosaccharides that shape colonization. Among infants who later developed NEC, consistent diet-linked patterns were less obvious, though when you compared formula-fed NEC infants to breast milk–fed controls, the formula group showed more Proteobacteria and less Firmicutes. Delivery didn’t produce stark case–control splits either. There wasn’t overall clustering by vaginal versus cesarean birth, but when you looked just at controls, cesarean delivery aligned with higher Firmicutes and vaginal birth with higher Bacteroidetes. So there are signals, but they braid together with disease status in complicated ways. If you’re hearing a theme, it’s this: the pre-NEC gut isn’t a chaos of random bugs, but the clarity of the signal depends on how and when you look. Pammi and colleagues took pains to standardize the “how”—uniform sequence processing, closed-reference OTU picking, strict quality thresholds, rarefaction to equalize depth—and then mapped the “when” to corrected gestational age rather than raw days of life. That choice matters because a twenty-six-week infant on day ten and a thirty-one-week infant on day ten aren’t at the same biological moment. The mean NEC diagnosis around thirty weeks corrected age anchors the timeline, and the microbial tilt toward Proteobacteria appears to gather momentum in the weeks just before. But they’re candid about the limits. The included studies spanned different hospitals and eras, with varying sample handling, DNA extraction kits, sequencing platforms, and treatment practices. Metadata weren’t always complete or harmonized across teams, which meant some analyses—like time-to-event models keyed to the exact day of NEC onset—weren’t feasible in the pooled data. Primer and region effects don’t vanish even with closed-reference matching; they can bias which taxa you detect and at what abundance. And while pooling boosts power, it can also wash out local signals. The upshot is both cautious and constructive: the directional pattern—more Proteobacteria and fewer Firmicutes and Bacteroidetes—precedes NEC, but the magnitude and the details shift with methods and exposures. Let’s pause on why that matters clinically. If you’re caring for a twenty-eight-week infant who’s wobbly on feeds and you see the stool microbiome lean into Proteobacteria, you might feel a twinge. The data say you’re not imagining it. But the same data say you can’t ignore antibiotics given last week, the 16S region your lab sequences, or whether this baby is on breast milk. Diversity by itself doesn’t give you a red light; composition and context do. The negative binomial model’s near miss on lower richness in NEC after adjusting for age—a p-value of zero point zero five three—captures the point. It’s suggestive, not sufficient. There’s also a methodological takeaway for microbiome science writ large. Choices that seem technical—the hypervariable region, the clustering algorithm, the reference database—aren’t background details. They shape conclusions. In this synthesis, V3 to V5 skews toward more Proteobacteria and fewer Firmicutes relative to V1 to V3, and studies themselves cluster by region and by laboratory when you project the UniFrac distances. If two papers disagree about whether Proteobacteria bloom before NEC, it may be that they looked through different keyholes. So where does this leave us? With a clearer, if still constrained, map. Before NEC, the gut community of preterm infants tilts toward Proteobacteria and away from Firmicutes and Bacteroidetes. That tilt tends to emerge in the late twenties of corrected gestational age and is shaped by antibiotics, feeding, and delivery. Alpha and beta diversity, the workhorse metrics of many microbiome papers, don’t cleanly separate cases from controls in this preclinical window. In other words, it’s the who more than the how many. Practically, that opens the door to smarter risk tracking, but only if we tighten the methods. Pammi and colleagues argue for standardization—common sampling windows aligned to corrected gestational age, harmonized DNA extraction and 16S protocols, and transparent reporting—so that future meta-analyses don’t have to fight the same headwinds. They also point toward richer data. Metagenomics to move beyond “who is there” to “what they can do,” and integration with host markers of inflammation to see whether microbial shifts and mucosal response rise together or in sequence. That’s how you get from association to mechanism. And that’s the bigger lesson of this work. In a condition as devastating as NEC, we don’t just want a single microbe to blame. We want a reliable early signal that stands up across hospitals, protocols, and patient stories. This synthesis doesn’t hand us a diagnostic test, but it does hand us a compass: watch for the Proteobacteria tilt, know the levers that amplify it, and build studies that reduce the noise. That’s how you turn thousands of tiny samples into something neonatologists can use at the bedside.

If you've ever set foot in a neonatal intensive care unit, you know it’s a place where hours stretch and every gram matters. In that world, necrotizing enterocolitis, or NEC, is the diagnosis everyone dreads. Among very low birthweight babies, those under 1500 grams, about seven percent will face it, and it accounts for up to five percent of neonatal intensive care admissions.

Mortality sits in a brutal band of approximately fifteen to thirty percent, and survivors can carry neurodevelopmental scars for years. That’s why a simple idea has had such staying power: maybe before NEC erupts, the gut ecosystem tips out of balance. Dysbiosis.

Fewer friendly neighbors, more potentially harmful ones. Animal studies hint at this—germ-free pups don’t get NEC—and human studies point in the same direction, but they’ve been small, scattered, and sometimes contradictory.

Pammi and colleagues decided to stop reading the tea leaves one cup at a time and brew a pot. Their question was sharp: can we define a microbial pattern that reliably precedes NEC in preterm infants, and can we explain why past studies disagree? They built a harmonized evidence base with tight study selection, uniform reprocessing of raw sequence data, and a plan to account for the usual suspects that muddy microbiome studies: antibiotics, feeding, delivery mode, and even which region of the 16S ribosomal RNA gene was sequenced. The goal wasn’t just to pool data. It was to compare apples with apples.

The search was broad and methodical. Using the Cochrane Neonatal Review Group strategy, they sifted through six thousand eight hundred and twelve records, screened four thousand two hundred and seventeen after de-duplication, and pulled twenty-four for full-text review. Fourteen met inclusion criteria.

Of those, nine teams shared not just summary results but the raw 16S reads plus clinical metadata. One study had to be set aside because the sequencing was too shallow. That left eight studies in the quantitative synthesis, involving one hundred and six infants who developed NEC, two hundred and seventy-eight controls, and a total of two thousand nine hundred and forty-four stool samples.

If you’re picturing a lot of tiny diapers and time points, you’re right—that’s the strength here. Many repeated samples per infant allow you to watch the microbiome evolve toward, or away from, disease.

To keep the analysis clean, every sequence ran through the same pipeline. They used Quantitative Insights Into Microbial Ecology, or QIIME, version one point eight point zero, which is a standard tool for microbial community analysis, and assigned operational taxonomic units, or OTUs, basically species bins, by matching reads to the GreenGenes reference. Only high-quality reads were kept: average quality scores had to be at least twenty-five, reads needed to be between two hundred and one thousand bases, barcodes and primers had to match exactly, and there couldn’t be ambiguous bases or long homopolymers.

Samples with fewer than one thousand two hundred reads were dropped, which is why that one study didn’t make the cut. Then they rarefied, essentially down-sampling so every sample had the same number of reads. That levels the playing field.

On timing, the diagnosis of NEC landed where clinicians would expect: a mean corrected gestational age—age adjusted for prematurity—of thirty point one weeks with a spread of about two point four weeks in the sixty-one infants where that date was available.

Here’s the headline. Before NEC shows itself clinically, the gut tilts in a consistent direction. Across corrected gestational ages, infants who went on to develop NEC had more Proteobacteria and fewer Firmicutes and Bacteroidetes compared with controls.

Think of Proteobacteria as a big tent that includes families like Enterobacteriaceae—Klebsiella and its cousins—and Firmicutes and Bacteroidetes as many of the usual early-life commensals. That tilt wasn’t random jitter; it began earlier and seemed to gather pace around twenty-seven weeks corrected age, approaching the thirty-week window when NEC typically hits. When Pammi and colleagues tested those phylum-level differences across all corrected ages, Proteobacteria rose and Firmicutes and Bacteroidetes fell with p-values below zero point zero five. The big picture is a directional shift, not a single smoking-gun microbe.

Now, you might ask: did overall diversity collapse before NEC? Surprisingly, not in a way you could rely on. When they measured alpha diversity—how many kinds of bacteria are present and how evenly they’re represented—using observed OTUs, the Shannon index, and the Simpson index, there weren’t consistent case–control differences across the data.

In a model tailored to count data, species richness did climb with gestational maturity—older guts harbor more kinds of microbes—with very strong support. After controlling for age, NEC cases trended toward lower richness than controls, but that signal just missed conventional significance. And when they looked at beta diversity—how different one sample is from another—using UniFrac, a distance that accounts for how related the bacteria are on the tree of life, the principal coordinates plots didn’t split into neat NEC versus control clouds.

Composition changed, but generic diversity metrics alone didn’t draw a diagnostic line.

Why so much fuzz? Methodology, for one. Not all 16S amplicons are created equal.

Three of the included studies sequenced the V1 to V3 hypervariable region of the 16S gene, while five targeted V3 to V5. Alpha diversity didn’t change much based on region—that is, counting species didn’t depend on which region you sequenced—but the relative proportions of phyla did. The V3 to V5 studies tended to report more Proteobacteria and fewer Firmicutes than V1 to V3 studies.

You can think of it as primer and region choice turning the lens slightly, making some groups look bigger and others smaller. In fact, even looking within controls, Shannon diversity differed when you compared V1 to V3 to V3 to V5, and that held when comparing NEC samples sequenced with V3 to V5 to controls sequenced with V1 to V3. Study-level quirks cropped up too.

Work by Normann, which used V3 to V5, and by Torraza, which used V1 to V3, leaned toward higher Firmicutes and lower Proteobacteria than the rest, reminding us that no pipeline erases all variability.

Clinical exposures also pulled hard on the microbiome. Antibiotics, unsurprisingly, left a deep footprint. In this meta-analysis, there were roughly one thousand eight hundred control samples collected during antibiotic exposure and just over one hundred collected without it; among NEC samples, about two hundred seventy were during antibiotics and a few dozen without.

Diversity shifted with antibiotics—OTU richness and the Shannon index differed significantly between exposed and unexposed controls—and the overall community structure clustered by exposure when you looked at UniFrac distances. Taxonomically, antibiotics pushed the gut toward Proteobacteria and away from Firmicutes, Actinobacteria, and Bacteroidetes. Zooming in to the genus level among controls, exposure was linked to higher abundances of Klebsiella, unclassified Enterobacteriaceae, Proteus, Paenibacillus, Epulopiscium, and Pseudomonas.

In antibiotic-free windows, Clostridium and unclassified Clostridiaceae were relatively more common. None of that is shocking if you’ve seen what broad-spectrum drugs do to a microbial community. It matters here because it can masquerade as, or amplify, the pre-NEC signature.

Feeding and delivery mode added nuance rather than clean separation. Looking within controls, formula-fed infants tended to have more Firmicutes, while those on breast milk showed higher relative Proteobacteria. It’s a reminder that human milk comes with its own microbes and oligosaccharides that shape colonization.

Among infants who later developed NEC, consistent diet-linked patterns were less obvious, though when you compared formula-fed NEC infants to breast milk–fed controls, the formula group showed more Proteobacteria and less Firmicutes. Delivery didn’t produce stark case–control splits either. There wasn’t overall clustering by vaginal versus cesarean birth, but when you looked just at controls, cesarean delivery aligned with higher Firmicutes and vaginal birth with higher Bacteroidetes.

So there are signals, but they braid together with disease status in complicated ways.

If you’re hearing a theme, it’s this: the pre-NEC gut isn’t a chaos of random bugs, but the clarity of the signal depends on how and when you look. Pammi and colleagues took pains to standardize the “how”—uniform sequence processing, closed-reference OTU picking, strict quality thresholds, rarefaction to equalize depth—and then mapped the “when” to corrected gestational age rather than raw days of life. That choice matters because a twenty-six-week infant on day ten and a thirty-one-week infant on day ten aren’t at the same biological moment.

The mean NEC diagnosis around thirty weeks corrected age anchors the timeline, and the microbial tilt toward Proteobacteria appears to gather momentum in the weeks just before.

But they’re candid about the limits. The included studies spanned different hospitals and eras, with varying sample handling, DNA extraction kits, sequencing platforms, and treatment practices. Metadata weren’t always complete or harmonized across teams, which meant some analyses—like time-to-event models keyed to the exact day of NEC onset—weren’t feasible in the pooled data.

Primer and region effects don’t vanish even with closed-reference matching; they can bias which taxa you detect and at what abundance. And while pooling boosts power, it can also wash out local signals. The upshot is both cautious and constructive: the directional pattern—more Proteobacteria and fewer Firmicutes and Bacteroidetes—precedes NEC, but the magnitude and the details shift with methods and exposures.

Let’s pause on why that matters clinically. If you’re caring for a twenty-eight-week infant who’s wobbly on feeds and you see the stool microbiome lean into Proteobacteria, you might feel a twinge. The data say you’re not imagining it.

But the same data say you can’t ignore antibiotics given last week, the 16S region your lab sequences, or whether this baby is on breast milk. Diversity by itself doesn’t give you a red light; composition and context do. The negative binomial model’s near miss on lower richness in NEC after adjusting for age—a p-value of zero point zero five three—captures the point. It’s suggestive, not sufficient.

There’s also a methodological takeaway for microbiome science writ large. Choices that seem technical—the hypervariable region, the clustering algorithm, the reference database—aren’t background details. They shape conclusions.

In this synthesis, V3 to V5 skews toward more Proteobacteria and fewer Firmicutes relative to V1 to V3, and studies themselves cluster by region and by laboratory when you project the UniFrac distances. If two papers disagree about whether Proteobacteria bloom before NEC, it may be that they looked through different keyholes.

So where does this leave us? With a clearer, if still constrained, map. Before NEC, the gut community of preterm infants tilts toward Proteobacteria and away from Firmicutes and Bacteroidetes.

That tilt tends to emerge in the late twenties of corrected gestational age and is shaped by antibiotics, feeding, and delivery. Alpha and beta diversity, the workhorse metrics of many microbiome papers, don’t cleanly separate cases from controls in this preclinical window. In other words, it’s the who more than the how many.

Practically, that opens the door to smarter risk tracking, but only if we tighten the methods. Pammi and colleagues argue for standardization—common sampling windows aligned to corrected gestational age, harmonized DNA extraction and 16S protocols, and transparent reporting—so that future meta-analyses don’t have to fight the same headwinds. They also point toward richer data.

Metagenomics to move beyond “who is there” to “what they can do,” and integration with host markers of inflammation to see whether microbial shifts and mucosal response rise together or in sequence. That’s how you get from association to mechanism.

And that’s the bigger lesson of this work. In a condition as devastating as NEC, we don’t just want a single microbe to blame. We want a reliable early signal that stands up across hospitals, protocols, and patient stories.

This synthesis doesn’t hand us a diagnostic test, but it does hand us a compass: watch for the Proteobacteria tilt, know the levers that amplify it, and build studies that reduce the noise. That’s how you turn thousands of tiny samples into something neonatologists can use at the bedside.

More in Nursing