Causes of variation in soil carbon simulations from CMIP5 Earth system models and comparison with observations

Katherine EO Todd-Brown, James T. Randerson, W. M. Post, Forrest M. Hoffman, C. Tarnocai, Edward A. G. Schuur, Steven AllisonView original
OverviewBalancedadam voice
If soil holds more carbon than all the world's plants and the atmosphere combined, then getting its behavior right in climate models matters enormously. If those models disagree with each other by a factor of six, then we don't have it right. Six. That's the spread in global soil carbon estimates across the eleven Earth system models that Todd-Brown and colleagues examined in their two thousand thirteen analysis — and that number is where this story begins. Soil organic carbon is the single largest carbon pool in the terrestrial biosphere. When soils warm, microbes break down that organic matter faster, releasing carbon dioxide back into the atmosphere — a feedback that could accelerate warming further. In high northern latitudes, permafrost soils are especially loaded with carbon accumulated over thousands of years, and they're particularly vulnerable to thaw. So, the ability of Earth system models — the large coupled climate-carbon models used in Intergovernmental Panel on Climate Change projections — to accurately represent existing soil carbon stocks is a prerequisite for trusting what those models say about the future. Todd-Brown and colleagues set out to test whether the Coupled Model Intercomparison Project phase five models, the generation of models used for the fifth Climate Model Intercomparison Project, actually meet that prerequisite. The answer, in short, is not really. The team compared soil carbon simulations from eleven independent modeling centers against two empirical databases: the Harmonized World Soil Database and the Northern Circumpolar Soil Carbon Database. Model estimates of global soil carbon stocks ranged from five hundred ten to three thousand forty petagrams of carbon — a petagram is roughly a trillion kilograms, an almost incomprehensible mass — while the Harmonized World Soil Database puts the real-world figure at around one thousand two hundred sixty petagrams of carbon, with a ninety-five percent confidence interval of eight hundred ninety to one thousand six hundred sixty petagrams. That five point nine-fold model spread doesn't just mean some models are a little high and some a little low. It means the range of answers spans more than the entire best-estimate value. The high-latitude story is worse. In the permafrost-rich northern regions, model estimates ranged from sixty to eight hundred twenty petagrams of carbon — more than a thirteenfold spread — against the Northern Circumpolar Soil Carbon Database estimate of about five hundred petagrams. These are the soils we're most worried about under warming, and the models disagree most severely about how much carbon they contain. So where does the disagreement come from? Todd-Brown and colleagues built a diagnostic tool: a reduced-complexity model. Instead of running the full machinery of an Earth system model, they asked whether a simple two-variable equation — soil carbon as a function of net primary productivity and soil temperature — could reproduce the spatial patterns that the full models produced. Net primary productivity is the amount of carbon that plants fix from the atmosphere each year, net of their own respiration; it's the rate at which organic material enters the soil system. In the reduced model, soil carbon at steady state equals net primary productivity divided by a decomposition rate constant, modified by a Q10 temperature sensitivity function — meaning decomposition speeds up exponentially with warming, referenced to a fifteen degrees Celsius baseline. The turnover time, which is simply one divided by that rate constant, tells you how long carbon sits in the soil before it's lost to decomposition. That simple model turned out to be a remarkably good mirror of the complex models. For nine of the eleven Earth system models, it explained between sixty-two and ninety-three percent of the grid-scale spatial variation in soil carbon. Complex models with many carbon pools, many processes, and decades of development are largely doing what a two-variable equation predicts. At the global scale, the reduced model explained ninety-eight percent of the variation in total soil carbon across all eleven models. When the team decomposed why models differ from each other, they found that differences in the decomposition parameterization alone — the rate constant and temperature sensitivity — explained sixty-four percent of the inter-model variance. Driving-variable differences, primarily net primary productivity, explained only about thirteen percent on their own, though substituting each model's net primary productivity while holding temperature constant still captured ninety-three percent of cross-model spread. The inferred turnover times ranged from about eleven years in the Community Climate System Model version four to thirty-seven years in the Model for Interdisciplinary Research on Climate Earth System Model — a three point six-fold range — and this variation in how fast models cycle carbon through the soil is a major lever on their wildly different global totals. Here's the twist. When the same reduced model was driven not by model outputs but by observed net primary productivity from the Moderate Resolution Imaging Spectroradiometer satellite data and real surface temperatures, it explained only ten percent of the spatial variation in actual Harmonized World Soil Database soil carbon. The same equation that mirrors what models do tells us almost nothing about where real-world soil carbon is concentrated. Something governs the geographic distribution of soil carbon in the real world that net primary productivity and temperature alone don't capture — something the models are also missing. Before concluding the models are simply wrong, though, Todd-Brown and colleagues complicate the picture in an important way: even the observations disagree with each other. In the high northern latitudes where the Harmonized World Soil Database and the Northern Circumpolar Soil Carbon Database overlap, the Pearson correlation between the two databases is just zero point thirty-three. A Pearson correlation of one is perfect agreement; zero means no relationship. So the two best empirical maps of northern soil carbon barely correspond spatially. This matters because the uncertainty isn't all on the model side. The "ground truth" is itself uncertain, and any effort to benchmark models is working against a fuzzy target. Scale is part of the story here too. At the biome level — grouping grid cells into broad ecosystem types like tropical forests or temperate grasslands — models do moderately well against observations, with R-squared values ranging from zero point thirty-eight to zero point ninety-seven, averaging zero point seventy-five. That sounds reasonable. But when comparisons are done at the one-degree grid scale, the spatial correlation between models and the Harmonized World Soil Database falls below zero point four for all eleven models, and root mean square errors run from nine point four to twenty point eight kilograms of carbon per square meter. Getting the biome totals roughly right while completely missing the local distribution is a problem, because regional feedbacks depend on where the carbon is, not just how much there is globally. So what would actually fix this? Todd-Brown and colleagues point to two concrete levers and a longer list of missing physics. The first lever is net primary productivity. Model skill at reproducing observed net primary productivity varied considerably — Pearson correlations against the Moderate Resolution Imaging Spectroradiometer ranged from zero point forty-eight to zero point seventy-five, while surface air temperature was much better constrained, at zero point ninety-three to zero point ninety-six. Since errors in net primary productivity flow directly into soil carbon, improving photosynthesis and autotrophic respiration algorithms would reduce propagated error downstream. The second lever is decomposition parameterization. The paper recommends using terrestrial radiocarbon measurements — both inventory data and vertical carbon-14 profiles through soil cores — to independently constrain turnover rates, rather than tuning decomposition parameters against the soil carbon totals they're meant to predict. Beyond those two levers, the paper flags a set of processes that current Earth system models largely omit: permafrost and thermokarst dynamics, peat accumulation and soil hydrology, cryoturbation, microbial and enzyme controls on decomposition, and the formation of chemically recalcitrant carbon compounds. These aren't minor additions. They're the mechanisms that likely explain why net primary productivity and temperature can't account for where soil carbon actually sits in the real world. What the sixfold spread in model estimates ultimately means is this: uncertainty about soil carbon stocks propagates forward into uncertainty about how much carbon could enter the atmosphere as those stocks warm and decompose. Soil carbon is not a passive reservoir. It's a potential source. And right now, our best models disagree about the size of that source by a factor of six. Todd-Brown and colleagues frame this not purely as a failure but as a diagnosis. We know that net primary productivity errors and decomposition parameterization drive the spread. We know that grid-scale patterns require processes beyond temperature and productivity. We know the observational databases themselves need better reconciliation. Knowing where the errors come from is real progress — because it turns an overwhelming problem into a set of tractable targets. The ground beneath our feet remains one of the least-understood variables in our climate future. Closing the gap between what models say and what the soil contains is some of the most consequential science being done right now. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

If soil holds more carbon than all the world's plants and the atmosphere combined, then getting its behavior right in climate models matters enormously. If those models disagree with each other by a factor of six, then we don't have it right. Six. That's the spread in global soil carbon estimates across the eleven Earth system models that Todd-Brown and colleagues examined in their two thousand thirteen analysis — and that number is where this story begins. Soil organic carbon is the single largest carbon pool in the terrestrial biosphere. When soils warm, microbes break down that organic matter faster, releasing carbon dioxide back into the atmosphere — a feedback that could accelerate warming further. In high northern latitudes, permafrost soils are especially loaded with carbon accumulated over thousands of years, and they're particularly vulnerable to thaw. So, the ability of Earth system models — the large coupled climate-carbon models used in Intergovernmental Panel on Climate Change projections — to accurately represent existing soil carbon stocks is a prerequisite for trusting what those models say about the future. Todd-Brown and colleagues set out to test whether the Coupled Model Intercomparison Project phase five models, the generation of models used for the fifth Climate Model Intercomparison Project, actually meet that prerequisite.

The answer, in short, is not really. The team compared soil carbon simulations from eleven independent modeling centers against two empirical databases: the Harmonized World Soil Database and the Northern Circumpolar Soil Carbon Database. Model estimates of global soil carbon stocks ranged from five hundred ten to three thousand forty petagrams of carbon — a petagram is roughly a trillion kilograms, an almost incomprehensible mass — while the Harmonized World Soil Database puts the real-world figure at around one thousand two hundred sixty petagrams of carbon, with a ninety-five percent confidence interval of eight hundred ninety to one thousand six hundred sixty petagrams. That five point nine-fold model spread doesn't just mean some models are a little high and some a little low. It means the range of answers spans more than the entire best-estimate value. The high-latitude story is worse. In the permafrost-rich northern regions, model estimates ranged from sixty to eight hundred twenty petagrams of carbon — more than a thirteenfold spread — against the Northern Circumpolar Soil Carbon Database estimate of about five hundred petagrams. These are the soils we're most worried about under warming, and the models disagree most severely about how much carbon they contain.

So where does the disagreement come from? Todd-Brown and colleagues built a diagnostic tool: a reduced-complexity model. Instead of running the full machinery of an Earth system model, they asked whether a simple two-variable equation — soil carbon as a function of net primary productivity and soil temperature — could reproduce the spatial patterns that the full models produced. Net primary productivity is the amount of carbon that plants fix from the atmosphere each year, net of their own respiration; it's the rate at which organic material enters the soil system. In the reduced model, soil carbon at steady state equals net primary productivity divided by a decomposition rate constant, modified by a Q10 temperature sensitivity function — meaning decomposition speeds up exponentially with warming, referenced to a fifteen degrees Celsius baseline. The turnover time, which is simply one divided by that rate constant, tells you how long carbon sits in the soil before it's lost to decomposition. That simple model turned out to be a remarkably good mirror of the complex models. For nine of the eleven Earth system models, it explained between sixty-two and ninety-three percent of the grid-scale spatial variation in soil carbon. Complex models with many carbon pools, many processes, and decades of development are largely doing what a two-variable equation predicts.

At the global scale, the reduced model explained ninety-eight percent of the variation in total soil carbon across all eleven models. When the team decomposed why models differ from each other, they found that differences in the decomposition parameterization alone — the rate constant and temperature sensitivity — explained sixty-four percent of the inter-model variance. Driving-variable differences, primarily net primary productivity, explained only about thirteen percent on their own, though substituting each model's net primary productivity while holding temperature constant still captured ninety-three percent of cross-model spread. The inferred turnover times ranged from about eleven years in the Community Climate System Model version four to thirty-seven years in the Model for Interdisciplinary Research on Climate Earth System Model — a three point six-fold range — and this variation in how fast models cycle carbon through the soil is a major lever on their wildly different global totals. Here's the twist. When the same reduced model was driven not by model outputs but by observed net primary productivity from the Moderate Resolution Imaging Spectroradiometer satellite data and real surface temperatures, it explained only ten percent of the spatial variation in actual Harmonized World Soil Database soil carbon. The same equation that mirrors what models do tells us almost nothing about where real-world soil carbon is concentrated.

Something governs the geographic distribution of soil carbon in the real world that net primary productivity and temperature alone don't capture — something the models are also missing. Before concluding the models are simply wrong, though, Todd-Brown and colleagues complicate the picture in an important way: even the observations disagree with each other. In the high northern latitudes where the Harmonized World Soil Database and the Northern Circumpolar Soil Carbon Database overlap, the Pearson correlation between the two databases is just zero point thirty-three. A Pearson correlation of one is perfect agreement; zero means no relationship. So the two best empirical maps of northern soil carbon barely correspond spatially. This matters because the uncertainty isn't all on the model side. The "ground truth" is itself uncertain, and any effort to benchmark models is working against a fuzzy target. Scale is part of the story here too. At the biome level — grouping grid cells into broad ecosystem types like tropical forests or temperate grasslands — models do moderately well against observations, with R-squared values ranging from zero point thirty-eight to zero point ninety-seven, averaging zero point seventy-five. That sounds reasonable.

But when comparisons are done at the one-degree grid scale, the spatial correlation between models and the Harmonized World Soil Database falls below zero point four for all eleven models, and root mean square errors run from nine point four to twenty point eight kilograms of carbon per square meter. Getting the biome totals roughly right while completely missing the local distribution is a problem, because regional feedbacks depend on where the carbon is, not just how much there is globally. So what would actually fix this? Todd-Brown and colleagues point to two concrete levers and a longer list of missing physics. The first lever is net primary productivity. Model skill at reproducing observed net primary productivity varied considerably — Pearson correlations against the Moderate Resolution Imaging Spectroradiometer ranged from zero point forty-eight to zero point seventy-five, while surface air temperature was much better constrained, at zero point ninety-three to zero point ninety-six. Since errors in net primary productivity flow directly into soil carbon, improving photosynthesis and autotrophic respiration algorithms would reduce propagated error downstream. The second lever is decomposition parameterization.

The paper recommends using terrestrial radiocarbon measurements — both inventory data and vertical carbon-14 profiles through soil cores — to independently constrain turnover rates, rather than tuning decomposition parameters against the soil carbon totals they're meant to predict. Beyond those two levers, the paper flags a set of processes that current Earth system models largely omit: permafrost and thermokarst dynamics, peat accumulation and soil hydrology, cryoturbation, microbial and enzyme controls on decomposition, and the formation of chemically recalcitrant carbon compounds. These aren't minor additions. They're the mechanisms that likely explain why net primary productivity and temperature can't account for where soil carbon actually sits in the real world. What the sixfold spread in model estimates ultimately means is this: uncertainty about soil carbon stocks propagates forward into uncertainty about how much carbon could enter the atmosphere as those stocks warm and decompose. Soil carbon is not a passive reservoir. It's a potential source. And right now, our best models disagree about the size of that source by a factor of six. Todd-Brown and colleagues frame this not purely as a failure but as a diagnosis. We know that net primary productivity errors and decomposition parameterization drive the spread.

We know that grid-scale patterns require processes beyond temperature and productivity. We know the observational databases themselves need better reconciliation. Knowing where the errors come from is real progress — because it turns an overwhelming problem into a set of tractable targets. The ground beneath our feet remains one of the least-understood variables in our climate future. Closing the gap between what models say and what the soil contains is some of the most consequential science being done right now. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Agricultural and Biological Sciences