Microbes as Engines of Ecosystem FunctionWhen Does Community Structure Enhance Predictions of Ecosystem Processes?
If microbes drive the carbon and nitrogen cycles that regulate Earth's climate, and if we can now sequence entire microbial communities from a handful of soil, then we should be able to predict how fast a forest floor respirates or how quickly a wetland converts nitrate to nitrogen gas, just by knowing who is living there. That reasoning is clean. It is also mostly wrong. Graham and colleagues tested this across eighty-two global datasets, and what they found is more useful than the clean version: microbes matter, but only under specific conditions, and knowing when those conditions apply is the real scientific advance. The study assembled eighty-two independently collected datasets that each measured site environmental conditions, microbial community data, and rates of carbon or nitrogen cycling, including respiration, nitrification, denitrification, and nitrogen mineralization. These are the processes that determine how quickly organic matter breaks down, how much nitrogen becomes available to plants, and how much ends up as atmospheric gas. Predicting them accurately feeds directly into climate models.
Graham and colleagues compared three predictor sets: environmental variables like pH, temperature, and moisture; microbial community structure, represented by diversity metrics like the Shannon index and ordination axes derived from multivariate taxonomic data; and microbial biomass, meaning microbial carbon or nitrogen content. By running all three sets through a multimodel inference framework, fitting every combination of variables within each predictor set and averaging across the best models, they could cleanly ask which information source adds value, and when. The headline result is unambiguous. Environmental variables were the strongest predictors of ecosystem process rates, averaging an adjusted R-squared of 0.56 across all eighty-two datasets. Microbial community structure alone averaged just 0.31. The gap between those two numbers, and its statistical significance, tells you that temperature, pH, and moisture do most of the heavy lifting. But here is the number that keeps the story interesting: environmental models left forty-four percent of variation unexplained on average. That is not a rounding error. That is nearly half the story missing.
Adding microbial community information to environmental models pushed the average adjusted R-squared from 0.56 to 0.65. This is statistically detectable, but not universal. By the study's criteria, improvement had to be both statistically significant and ecologically meaningful, defined as an increase in adjusted R-squared greater than ten percent of the environmental model's value. Only twenty-nine percent of datasets were significantly improved by adding microbial data, with an average gain of 0.08. Let that land: most of the time, once you know the environment, knowing the microbes doesn't move the needle. But that unresolved forty-four percent is exactly the space the paper then spends its energy mapping. Two distinct patterns emerged about when microbial data does help, and they point in different directions depending on the process. The first pattern involves processes carried out by phylogenetically narrow guilds—small, specialized groups of organisms that are the only ones capable of a given reaction. Nitrification is the textbook case. It is performed almost exclusively by ammonia-oxidizing bacteria and archaea, and it was the clearest success story for microbial data. Across fourteen nitrification datasets, models built on functional gene abundance, which are direct counts of the genes encoding nitrification enzymes, averaged an adjusted R-squared of 0.61. Models built on community diversity metrics averaged just 0.21.
That difference was statistically significant, and half of the environmental models for nitrification were improved by adding functional gene data. The message is direct: if only a narrow slice of the microbial world does the job, counting the genes those organisms carry tells you something that temperature and pH cannot. The second pattern runs in the opposite direction. For facultative processes—reactions that many different organisms can perform depending on conditions—community diversity metrics added more predictive value than functional gene counts. Denitrification, which is carried out by diverse facultative anaerobes, followed this logic. In three of four denitrification datasets where sixteen S ribosomal ribonucleic acid gene diversity was added to environmental models, models improved by an average adjusted R-squared increase of 0.13, whereas only three of eleven models improved with functional gene abundance, with an average gain of 0.04. When the metabolic function is spread across many lineages, knowing the diversity of the community tells you more than counting any single functional marker. Respiration added a third wrinkle. Environmental variables explained respiration well on their own, averaging an adjusted R-squared of 0.66 across twenty-six datasets. Microbial-only models averaged just 0.29.
However, the combination of biomass and community structure was notably more powerful than either alone: combining both improved fifty-three percent of respiration models with an average adjusted R-squared increase of 0.15, compared to thirty-five percent improved by microbial biomass alone, with an average gain of 0.09. The reason is informative—biomass and community structure were only weakly correlated with each other, averaging an adjusted R-squared of just 0.11 in redundancy analyses, meaning they capture different aspects of the microbial system. When you use both, you get additive signal. So why does microbial community data fail to add predictive power most of the time? Graham and colleagues identify several mechanisms, and together they form a coherent account. The first is dormancy. DNA-based community surveys detect every organism present, active or not. But the percentage of a soil community actually catalyzing reactions at any moment can vary enormously with resource availability and disturbance history. You are counting players on the bench alongside players on the field, with no way to distinguish them. The second mechanism is functional redundancy. Many different species can perform the same metabolic reaction, so which species are present matters less than whether the function is represented at all. This is why broad, facultative processes resist prediction from community composition — the specific cast of characters is largely interchangeable.
The third mechanism is a scale mismatch. Process-rate measurements integrate over time and space in ways that community measurements do not. An environmental variable like soil pH is measured at the plot scale, while a DNA extraction comes from a fraction of a gram. They are not describing the same thing at the same resolution. Fourth, and perhaps most quietly important: microbial community composition is often itself a reflection of the environment. Gene-abundance data correlated with environmental variables at an average adjusted R-squared of 0.36, while community diversity metrics correlated with environmental variables at just 0.20. When both community data and environmental data are telling the same story, adding one to a model that already contains the other adds little new information. These are not excuses for a failed analysis. They are the actual scientific output. Understanding why the link between community structure and process rate is so often weak tells you exactly what to measure next. Graham and colleagues close with a specific argument about how to close the forty-four percent gap. For narrow, obligate processes, functional gene abundance is the right tool, as demonstrated by the nitrification results. For broad, facultative processes, diversity- and trait-aware approaches are more appropriate.
Combining community structure with biomass measurements captures additive information that neither provides alone, as the respiration results show. And across all processes, reducing dormancy artifacts, matching spatial and temporal scales between community and process measurements, and incorporating transcriptomic or proteomic data—which can distinguish active from dormant organisms—are the next methodological steps. Trait-based targeting and cross-scale modeling, the paper argues, may be the path to making microbial community data genuinely predictive rather than occasionally informative. What this study ultimately provides is a map of conditions. Narrow guilds, functional genes, and obligate processes: microbes predict. Broad metabolisms, facultative reactions, and high functional redundancy: diversity metrics matter more, but gains are modest. Biomass plus community structure together is better than either alone for respiration. And in all cases, the environment sets the baseline that microbial data has to beat. That is not a failure of microbial ecology. It is its current frontier—a forty-four percent gap, precisely located, with specific tools identified to close it. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Transport Distance of Invertebrate Environmental DNA in a Natural River
- Moving in the Anthropocene: Global reductions in terrestrial mammalian movements
- A Novel Coronavirus Genome Identified in a Cluster of Pneumonia Cases — Wuhan, China 2019−2020
- Vapors Produced by Electronic Cigarettes and E-Juices with Flavorings Induce Toxicity, Oxidative Stress, and Inflammatory Response in Lung Epithelial Cells and in Mouse Lung
- Differential climate impacts for policy-relevant limits to global warming: the case of 1.5 °C and 2 °C
- Absorption Angstrom Exponent in AERONET and related data as an indicator of aerosol composition