The CAMELS data setcatchment attributes and meteorology for large-sample studies
Picture a hydrologist staring at a model that keeps failing. Not because the physics are wrong, but because nobody has consistent, reliable data on what the watershed actually looks like — its soils, its geology, how much of its precipitation falls as snow, and how quickly it drains after a storm. That's the problem Addor and colleagues set out to solve. The result is CAMELS — Catchment Attributes and Meteorology for Large-sample Studies — covering six hundred seventy-one catchments across the contiguous United States, six classes of attributes, freely available to anyone who needs it. Catchment attributes are the descriptors of a landscape that together determine how a watershed stores and moves water. Climate, soils, geology, land cover, and topography — these controls interact in nonlinear ways. This means understanding how a watershed responds to rainfall requires large, consistent samples, not a handful of isolated gauges. Before CAMELS, the most widely used continental-scale compilation was MOPEX — the Model Parameter Estimation Experiment — which contained four hundred thirty-eight catchments concentrated mainly in the eastern half of the country, leaving the Rocky Mountains severely underrepresented. MOPEX relied on older point observations, lacked geology data entirely, and only fifty-two of its catchments overlap with those in CAMELS. Addor and colleagues were explicit: this wasn't just about adding more catchments.
It was about broader geographic balance, newer underlying data, and — critically — honest documentation of where the data fall short. CAMELS pairs two complementary components. The hydrometeorological backbone comes from Newman and colleagues, who assembled daily meteorological forcings from three gridded products — Daymet, NLDAS, and Maurer — along with United States Geological Survey streamflow records for six hundred seventy-one minimally disturbed catchments, each with at least twenty years of continuous discharge spanning from nineteen eighty to two thousand fifteen. Addor and colleagues then computed a harmonized set of catchment attributes on top of that foundation, organized into six classes: topography, climate, streamflow signatures, land cover, soil, and geology. Topography is the most straightforward — mean elevation, mean slope, and catchment area derived from digital elevation models. Five catchments exceed ten thousand square kilometers, four of them in the Great Plains, and the authors note that catchment averages become less meaningful as the area grows. Climate indices were computed from the Daymet forcing over the period from October nineteen eighty-nine through September two thousand nine — a total of twenty hydrological years — and summarize mean precipitation and potential evapotranspiration, aridity, snow fraction, precipitation seasonality, and the frequency and duration of dry days and high-precipitation events.
Before any results, the method for catchment delineation deserves a note: Addor and colleagues compared two polygon sources, the geospatial fabric used by Newman and colleagues and the GAGES II dataset. Eight catchments had an absolute relative area error greater than one hundred percent between the two sources, and sixty-two had errors greater than ten percent — a real warning for modelers who assume catchment boundaries are unambiguous. The climate results reveal strong, spatially coherent gradients. Aridity — the ratio of mean annual potential evapotranspiration to mean annual precipitation — is the organizing variable. Where aridity exceeds one, evaporative demand outstrips supply; where it falls below one, precipitation wins. Parts of the Great Plains sit well above one. Much of the Pacific Northwest sits below 0.5, meaning precipitation is roughly double potential evapotranspiration. Snow fraction, computed from daily temperatures using a zero degrees Celsius threshold, follows elevation and latitude, peaking in the Rockies. Precipitation seasonality, summarized by a sine-curve timing metric, captures the split between winter-dominated western catchments and summer-dominated central ones.
These climatic fingerprints map directly onto streamflow behavior. The runoff ratio — mean daily discharge divided by mean daily precipitation — collapses in arid regions. In the Great Plains, more than eighty percent of precipitation evaporates, pushing runoff ratios below 0.2 and mean annual discharges down to around 0.3 millimeters per day. In the Pacific Northwest, where aridity is below 0.5, both runoff ratio and mean annual discharge are substantially higher, and most streamflow arrives in the first half of the year. The baseflow index — the fraction of total discharge attributable to slow, sustained drainage rather than direct runoff — and the slope of the flow duration curve work in tandem. A steeper curve combined with a low baseflow index signals a flashy catchment. That combination appears in a band from eastern Kansas to Kentucky, where rapid runoff dominates. Rocky Mountain catchments store water as snow and release it late, while a swath from eastern Texas to South Carolina shows early seasonal flows. At the extreme end, the ninety-fifth percentile flow, or Q95, is more than ten times higher in the Pacific Northwest and Appalachians than in the most arid catchments. In those arid basins, runoff is also more sensitive to year-to-year precipitation anomalies: when rain is scarce, small deviations have outsized effects on what actually reaches the stream.
Land cover, soil, and geology required the most methodological ingenuity — and carry the most uncertainty. For vegetation, CAMELS used one-kilometer MODIS products from twenty-oh-two to twenty-fourteen to compute two indicators: leaf area index, or LAI, which measures one-sided green leaf area per unit ground area, and green vegetation fraction, or GVF, which captures the fraction of a grid cell covered by vegetation. The team extracted maximum monthly LAI and the seasonal amplitude — the swing between maximum and minimum — to characterize both density and seasonality. LAI and GVF correlate closely, but the authors flag a hard limit: MODIS cannot reliably separate vertical leaf stacking from horizontal canopy packing. Both metrics capture overall vegetation amount. They cannot distinguish a dense canopy from a sprawling one. Soil attributes were built primarily from STATSGO — the State Soil Geographic database — processed by Miller and White, who discretized the top 2.5 meters into eleven layers of increasing thickness. Saturated hydraulic conductivity and porosity per layer were estimated using the regressions of Cosby and colleagues. At the catchment scale, CAMELS computes thickness-weighted averages over the top 1.5 meters — summing layer value times layer thickness, then dividing by cumulative depth — and uses the harmonic mean for hydraulic conductivity.
But the deeper you go, the thinner the data gets. Only about 2.5 percent of STATSGO components have layers extending below two hundred three centimeters. Roughly half report identical minimum and maximum depth to bedrock of one hundred fifty-two centimeters, which means bedrock was never actually reached — the record just bottomed out. Comparing STATSGO to independent depth-to-bedrock estimates from Pelletier and colleagues, forty-seven percent of CAMELS catchments have bedrock deeper than 1.5 meters, beyond STATSGO's reach entirely. Twenty-four percent have bedrock deeper than fifteen meters — an order of magnitude beyond coverage. The authors restricted their soil analyses to the top 1.5 meters because of this, and they are direct about what that means for anyone trying to model deep percolation or slow groundwater recharge. Geology was characterized using two global datasets. GLiM — the Global Lithological Map of Hartmann and Moosdorf — contains roughly one point two million polygons and classifies rock types into sixteen first-level classes; CAMELS records area fractions per class for each catchment.
GLHYMPS, from Gleeson and colleagues, maps porosity and permeability derived from those lithologies. CAMELS computes catchment-averaged porosity using the arithmetic mean and permeability using the geometric mean, reflecting the different statistical behavior of those two properties. Across the dataset, siliciclastic sedimentary rocks appear in thirty-four percent of catchments, unconsolidated sediments in nineteen percent, metamorphic rocks in sixteen percent, and carbonate sedimentary rocks in twelve percent. Eighteen percent of catchments contain only a single lithological type; in eleven percent, the dominant class covers less than half the area. One important caveat: both GLiM and GLHYMPS show unrealistic spatial discontinuities at jurisdictional boundaries — the North and South Dakota region is one example — where underlying data sources change abruptly. Addor and colleagues treat these limitations not as failures but as documentation. That transparency is a design choice. Users can read the metadata, understand which attributes carry the most uncertainty, and make informed decisions about which catchments to include in a given analysis. That is precisely what was missing from earlier compilations.
CAMELS is freely available, and together the hydrometeorological series and the catchment attribute tables enable questions that simply weren't tractable before. Does a soil hydraulic conductivity parameter estimated from textural fractions actually predict runoff ratios across hundreds of basins? Does a model calibrated in humid catchments hold up in arid ones? Does baseflow behavior track with geology as consistently as aridity tracks with climate? With six hundred seventy-one catchments spanning the full range of conditions across the contiguous United States, those questions now have a testing ground. Planned extensions include drainage density, stream-order statistics, and more refined uncertainty characterization for forcing and streamflow. The modeler who once lacked consistent data across many basins now has an open, documented foundation — one that tells you not just what the numbers are, but where to be careful with them. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Assessing “Dangerous Climate Change”: Required Reduction of Carbon Emissions to Protect Young People, Future Generations and Nature
- Plastics Derived Endocrine Disruptors (BPA, DEHP and DBP) Induce Epigenetic Transgenerational Inheritance of Obesity, Reproductive Disease and Sperm Epimutations
- Community genomic analyses constrain the distribution of metabolic traits across the Chloroflexi phylum and indicate roles in sediment carbon cycling
- The Impacts of Oil Palm on Recent Deforestation and Biodiversity Loss
- Transport Distance of Invertebrate Environmental DNA in a Natural River
- Microbes as Engines of Ecosystem Function: When Does Community Structure Enhance Predictions of Ecosystem Processes?