Analysis of 100,000 human cancer genomes reveals the landscape of tumor mutational burden

Zachary R. Chalmers, Caitlin Connelly, David Fabrizio, Laurie M. Gay, Siraj M. Ali, Riley Ennis, Alexa B. Schrock, Brittany Campbell, Adam Shlien, Juliann Chmielecki, Franklin W. Huang, Yuting He, James Sun, Uri Tabori, Mark Kennedy, Daniel S. Lieber, Steven Roels, Jared White, Geoffrey A. Otto, Jeffrey S. Ross, Levi A. Garraway, Vincent A. Miller, P. J. Stephens, Garrett M. FramptonView original
OverviewBalancedalloy voice
If you’ve followed cancer immunotherapy over the last decade, you’ve heard two acronyms tossed around like passwords: programmed death-ligand 1, or PD-L1, and tumor mutational burden, or TMB. PD-L1 is the protein we stain for on tumor cells to guess whether a checkpoint drug might work. Tumor mutational burden is the rough count of coding mutations a tumor has piled up, serving as a proxy for how many neoantigens it might display to the immune system. Clinicians like the idea of TMB because, unlike an antibody stain, it’s a number you can compute and compare. But here’s the practical issue: how do you measure it well enough, fast enough, and affordably enough for real-world clinics? That’s where Chalmers, Connelly, Frampton, and colleagues stepped in. They built a targeted comprehensive genomic profiling assay, which you can think of as a smartly chosen slice of the genome that spans about 1.1 megabases across roughly 315 cancer genes. They sequence those exons deeply, usually with coverage greater than 500-fold, and they made sure critical DNA repair genes like the PMS2 gene were captured across panel versions. On that slice, they count somatic, coding substitutions and small insertions or deletions. They even keep synonymous changes, not because those generate neoantigens, but because they stabilize the sampling and provide a truer read on the underlying mutation rate. Since there’s no matched normal for most clinical cases, they rely on rigorous germline filtering. This includes algorithms for somatic-germline zygosity, recurrent germline flags, and known germline data from the Single Nucleotide Polymorphism database, or dbSNP, and the Exome Aggregation Consortium database, to eliminate inherited variants and lab artifacts. Clean up the calls, divide by the megabases interrogated, and you’ve got TMB. Because microsatellite instability, or MSI, often travels with hypermutation, they incorporated MSI detection into the same profiling run. Here’s the trick they used: pick one hundred fourteen homopolymer loci in introns, measure the length variation at each locus in both directions, and you get two hundred twenty-eight data points per tumor. Project those onto a single axis through principal components analysis, and then, using training data and unsupervised clustering, classify each case as microsatellite-stable, MSI-high, or ambiguous. All this is done from the same sequencing library. In practice, these were clinical specimens with at least twenty percent tumor and enough DNA to create a robust library, so the assay could run at scale without needing an overwhelming amount of tissue. The obvious validation question is, does this targeted panel really reflect what whole-exome sequencing would reveal? In head-to-head tests on tumors that had both, the correlation between panel TMB and exome TMB was R squared of 0.74 across twenty-nine cases. It’s not perfect, but surprisingly tight given that the panel is just one twentieth of the exome. Precision was even cleaner. When they resequenced sixty tumors on the panel, the TMB values practically aligned along a line, with R squared of 0.98. To stress-test the design itself, they ran simulations: starting with tumors at three different mutation rates—think ten, twenty, and one hundred mutations per megabase—then simulating that they sequenced anywhere from a sliver of 0.2 megabases to a broad ten megabases, repeating that a thousand times and assessing how far off they would be. As you shrink the sequenced region, the variance increases, especially for low and mid TMB. Below roughly half a megabase, error balloons; around 1.1 megabases, sampling noise settles to something clinicians can work with. They even analyzed nearly nine thousand exomes from The Cancer Genome Atlas and checked how counts in their three hundred fifteen genes matched with exome-wide counts. The direction and strength matched closely, with an R squared around 0.98 for that relationship, reinforcing the point that a well-chosen slice can effectively represent the whole. With that groundwork laid, they aimed the panel at the real world. More than one hundred thousand tumors came through the door. After removing duplicates and samples with insufficient coverage, they analyzed ninety-two thousand four hundred thirty-nine cancers—representing hundreds of histologies, people across the age spectrum, and all the complexities of clinical practice. It’s the scale needed to draw a clean map of mutational burden across cancer, and the map is broad. Overall, the median TMB was three point six mutations per megabase, with an astonishing range from zero to more than one thousand. There wasn’t a notable difference between men and women in that median number, which is helpful—it indicates that sex isn’t the first factor to consider when evaluating mutational load. Age, however, did matter. One could almost see the clock ticking in the data. Using a simple linear model, TMB increased about two and a half times between childhood and late life. The median was around one point seven mutations per megabase at age ten, climbing to about four point five by the late eighties. This makes intuitive sense—mutations accumulate over time—but it’s useful to have the slope. It provides clinicians with a baseline expectation and helps explain why pediatric tumors, overall, appear quieter than adult ones. The disease-by-disease view adds texture. Among the one hundred sixty-seven tumor types with enough samples to be confident, medians ranged from under one mutation per megabase in bone marrow myelodysplastic syndromes to roughly forty-five in skin squamous cell carcinoma. That’s a fifty-fold swing between typical cases. And here’s the crucial clinical point: high TMB appears in pockets across many diseases, not just the usual suspects. Defining “high” the same way the field often does—more than twenty mutations per megabase—you observe at least ten percent high-TMB cases in twenty different tumor types, spanning eight organ systems. Stretch that bar to five percent, and thirty-eight tumor types qualify across nineteen tissues. Even in areas like soft tissue angiosarcoma, where the median is around three point eight, more than one in eight patients crossed that high-TMB threshold. This highlights the message: don’t dismiss rarer tumors. There are immunotherapy-relevant outliers lurking in many of them. Since MSI is the prototypical path to hypermutation in certain cancers, the team examined how it overlapped with TMB across sixty-two thousand one hundred fifty cases where an MSI call was possible. The overlap was strong, but not absolute. Most MSI-high tumors were also high TMB—about eighty-three percent—but the reverse wasn’t true. Only sixteen percent of high-TMB tumors were MSI-high. Nearly all MSI-high tumors, ninety-seven percent, at least exceeded ten mutations per megabase, but high TMB often appeared without MSI, particularly in melanoma, squamous skin cancers, and lung cancers. In gastrointestinal cancers, like stomach and small intestine, MSI and high TMB tended to coincide. This means that MSI and TMB measure related, but not identical, forms of genomic instability, and the tissue context indicates which mechanism is causing the damage. What about the genes associated with high TMB? When sifting through the dataset, certain names keep appearing. A total of two hundred fifty-seven genes had somatic alterations that corresponded with increased TMB at very stringent false discovery thresholds, and forty-eight of those had large, consistent effects based on factor loading. The usual suspects in mismatch repair lit up—the MSH2 gene, the MSH6 gene, the MLH1 gene, the PMS2 gene—along with the proofreading DNA polymerase POLE. When those pathways were combined, functional mutations in mismatch repair or POLE collectively appeared in roughly thirteen point five percent of cases. This represents a significant portion of the hypermutated landscape explained by repair failures or proofreading errors. Then there’s an unexpected twist: a recurrent hit in the promoter of the PMS2 gene, rather than in the coding sequence. Frampton’s team identified a cluster of twelve positions in the PMS2 promoter that were repeatedly mutated in skin cancers and were strongly associated with high TMB. One of those positions, at chromosome seven position six million forty-eight thousand seven hundred eighty-eight, showed an almost comically small p-value when linked to TMB in melanoma. The effect size is striking. In melanoma, having a mutation in the PMS2 promoter resulted in a more than five-fold increase in median TMB compared to tumors with the wild-type promoter. These aren’t just one-off anomalies; about ten percent of melanomas carried at least one of these promoter mutations. The frequency was even higher in basal cell carcinoma, roughly one in four, and close to one in five in skin squamous cell carcinoma. Other tumor types had lower rates, suggesting a mutagenesis process—like ultraviolet light—that’s particularly effective at causing these promoter changes in skin. Were these inherited? That was the next check. Most promoter variants that could be classified—about ninety-three percent—appeared somatic based on their somatic-germline algorithm, and their allele fractions hovered around twenty-five percent, consistent with subclonal acquisition. To validate this finding in an independent set, they examined The Cancer Genome Atlas melanoma cohort. The same promoter positions appeared there too. Four out of fifty melanomas in that set carried them, around eight percent, and when the analysis was expanded, the overall estimate returned to roughly ten percent with reasonable confidence bounds. Functionally, several computational predictors pointed to two promoter coordinates—six million forty-eight thousand seven hundred sixty and six million forty-eight thousand eight hundred twenty-four—as the most likely to alter gene regulation. When they investigated whether these promoter mutations tended to co-occur with other known drivers in melanoma, none held up after adjusting for TMB. This indicates a clean, recurrent regulatory hotspot linked to mutation load in skin cancers, acting predominantly on its own. Stepping back, a few themes emerge. First, a thoughtfully designed panel of about 1.1 megabases can substitute for exome sequencing when the goal is to quantify tumor mutational burden. The evidence is solid: there is a strong correlation against exome in paired tumors, nearly perfect reproducibility on replicates, and simulation studies that clarify why panel size matters—significantly shrinking below half a megabase leads to noisy results. Second, TMB is a landscape, not a single peak. The median is modest across cancers, but high-TMB clusters emerge in dozens of diseases, including rare ones, and they become more common as patients age. Third, while MSI captures part of the hypermutation story, particularly in gastrointestinal contexts, many high-TMB tumors aren’t MSI-high; other mechanisms—including alterations in mismatch repair genes, POLE mutations, and regulatory hits like the PMS2 promoter mutations—are involved. There are limits to this approach. As Frampton and colleagues clarify, a targeted panel isn’t the best tool for discovering every neoantigen; by its nature, it only observes a sliver of the genome and won’t capture the entire antigenic catalog. Germline filtering, although stringent, remains an area to refine further, especially for ancestries that have been underrepresented in reference databases. But for the fundamental, often repeated question in clinics—does this tumor carry enough mutational weight that a checkpoint inhibitor might have a chance?—this approach provides an answer that aligns well with whole-exome truth and can be scaled effectively. Clinically, this means TMB joins MSI, PD-L1, and specific drivers as one of the key pillars to evaluate. It’s not a kingmaker on its own, but it helps identify patients across a variety of tumor types who might benefit from immunotherapy. It also encourages us to examine genomes with both breadth and nuance: the base-by-base counts, the repair pathways that have failed, and the peculiar regulatory lesion that tips a balance. As Chalmers and colleagues demonstrated by mapping more than one hundred thousand cancers, when you measure the right factors in the right way, the signal begins to stand out. You obtain a map you can utilize. And in oncology, that map often represents the difference between guessing and making informed choices.

If you’ve followed cancer immunotherapy over the last decade, you’ve heard two acronyms tossed around like passwords: programmed death-ligand 1, or PD-L1, and tumor mutational burden, or TMB. PD-L1 is the protein we stain for on tumor cells to guess whether a checkpoint drug might work. Tumor mutational burden is the rough count of coding mutations a tumor has piled up, serving as a proxy for how many neoantigens it might display to the immune system.

Clinicians like the idea of TMB because, unlike an antibody stain, it’s a number you can compute and compare. But here’s the practical issue: how do you measure it well enough, fast enough, and affordably enough for real-world clinics?

That’s where Chalmers, Connelly, Frampton, and colleagues stepped in. They built a targeted comprehensive genomic profiling assay, which you can think of as a smartly chosen slice of the genome that spans about 1.1 megabases across roughly 315 cancer genes. They sequence those exons deeply, usually with coverage greater than 500-fold, and they made sure critical DNA repair genes like the PMS2 gene were captured across panel versions.

On that slice, they count somatic, coding substitutions and small insertions or deletions. They even keep synonymous changes, not because those generate neoantigens, but because they stabilize the sampling and provide a truer read on the underlying mutation rate. Since there’s no matched normal for most clinical cases, they rely on rigorous germline filtering.

This includes algorithms for somatic-germline zygosity, recurrent germline flags, and known germline data from the Single Nucleotide Polymorphism database, or dbSNP, and the Exome Aggregation Consortium database, to eliminate inherited variants and lab artifacts. Clean up the calls, divide by the megabases interrogated, and you’ve got TMB.

Because microsatellite instability, or MSI, often travels with hypermutation, they incorporated MSI detection into the same profiling run. Here’s the trick they used: pick one hundred fourteen homopolymer loci in introns, measure the length variation at each locus in both directions, and you get two hundred twenty-eight data points per tumor. Project those onto a single axis through principal components analysis, and then, using training data and unsupervised clustering, classify each case as microsatellite-stable, MSI-high, or ambiguous.

All this is done from the same sequencing library. In practice, these were clinical specimens with at least twenty percent tumor and enough DNA to create a robust library, so the assay could run at scale without needing an overwhelming amount of tissue.

The obvious validation question is, does this targeted panel really reflect what whole-exome sequencing would reveal? In head-to-head tests on tumors that had both, the correlation between panel TMB and exome TMB was R squared of 0.74 across twenty-nine cases. It’s not perfect, but surprisingly tight given that the panel is just one twentieth of the exome.

Precision was even cleaner. When they resequenced sixty tumors on the panel, the TMB values practically aligned along a line, with R squared of 0.98. To stress-test the design itself, they ran simulations: starting with tumors at three different mutation rates—think ten, twenty, and one hundred mutations per megabase—then simulating that they sequenced anywhere from a sliver of 0.2 megabases to a broad ten megabases, repeating that a thousand times and assessing how far off they would be.

As you shrink the sequenced region, the variance increases, especially for low and mid TMB. Below roughly half a megabase, error balloons; around 1.1 megabases, sampling noise settles to something clinicians can work with. They even analyzed nearly nine thousand exomes from The Cancer Genome Atlas and checked how counts in their three hundred fifteen genes matched with exome-wide counts.

The direction and strength matched closely, with an R squared around 0.98 for that relationship, reinforcing the point that a well-chosen slice can effectively represent the whole.

With that groundwork laid, they aimed the panel at the real world. More than one hundred thousand tumors came through the door. After removing duplicates and samples with insufficient coverage, they analyzed ninety-two thousand four hundred thirty-nine cancers—representing hundreds of histologies, people across the age spectrum, and all the complexities of clinical practice.

It’s the scale needed to draw a clean map of mutational burden across cancer, and the map is broad. Overall, the median TMB was three point six mutations per megabase, with an astonishing range from zero to more than one thousand. There wasn’t a notable difference between men and women in that median number, which is helpful—it indicates that sex isn’t the first factor to consider when evaluating mutational load.

Age, however, did matter. One could almost see the clock ticking in the data. Using a simple linear model, TMB increased about two and a half times between childhood and late life.

The median was around one point seven mutations per megabase at age ten, climbing to about four point five by the late eighties. This makes intuitive sense—mutations accumulate over time—but it’s useful to have the slope. It provides clinicians with a baseline expectation and helps explain why pediatric tumors, overall, appear quieter than adult ones.

The disease-by-disease view adds texture. Among the one hundred sixty-seven tumor types with enough samples to be confident, medians ranged from under one mutation per megabase in bone marrow myelodysplastic syndromes to roughly forty-five in skin squamous cell carcinoma. That’s a fifty-fold swing between typical cases.

And here’s the crucial clinical point: high TMB appears in pockets across many diseases, not just the usual suspects. Defining “high” the same way the field often does—more than twenty mutations per megabase—you observe at least ten percent high-TMB cases in twenty different tumor types, spanning eight organ systems. Stretch that bar to five percent, and thirty-eight tumor types qualify across nineteen tissues.

Even in areas like soft tissue angiosarcoma, where the median is around three point eight, more than one in eight patients crossed that high-TMB threshold. This highlights the message: don’t dismiss rarer tumors. There are immunotherapy-relevant outliers lurking in many of them.

Since MSI is the prototypical path to hypermutation in certain cancers, the team examined how it overlapped with TMB across sixty-two thousand one hundred fifty cases where an MSI call was possible. The overlap was strong, but not absolute. Most MSI-high tumors were also high TMB—about eighty-three percent—but the reverse wasn’t true.

Only sixteen percent of high-TMB tumors were MSI-high. Nearly all MSI-high tumors, ninety-seven percent, at least exceeded ten mutations per megabase, but high TMB often appeared without MSI, particularly in melanoma, squamous skin cancers, and lung cancers. In gastrointestinal cancers, like stomach and small intestine, MSI and high TMB tended to coincide.

This means that MSI and TMB measure related, but not identical, forms of genomic instability, and the tissue context indicates which mechanism is causing the damage.

What about the genes associated with high TMB? When sifting through the dataset, certain names keep appearing. A total of two hundred fifty-seven genes had somatic alterations that corresponded with increased TMB at very stringent false discovery thresholds, and forty-eight of those had large, consistent effects based on factor loading.

The usual suspects in mismatch repair lit up—the MSH2 gene, the MSH6 gene, the MLH1 gene, the PMS2 gene—along with the proofreading DNA polymerase POLE. When those pathways were combined, functional mutations in mismatch repair or POLE collectively appeared in roughly thirteen point five percent of cases. This represents a significant portion of the hypermutated landscape explained by repair failures or proofreading errors.

Then there’s an unexpected twist: a recurrent hit in the promoter of the PMS2 gene, rather than in the coding sequence. Frampton’s team identified a cluster of twelve positions in the PMS2 promoter that were repeatedly mutated in skin cancers and were strongly associated with high TMB. One of those positions, at chromosome seven position six million forty-eight thousand seven hundred eighty-eight, showed an almost comically small p-value when linked to TMB in melanoma.

The effect size is striking. In melanoma, having a mutation in the PMS2 promoter resulted in a more than five-fold increase in median TMB compared to tumors with the wild-type promoter. These aren’t just one-off anomalies; about ten percent of melanomas carried at least one of these promoter mutations.

The frequency was even higher in basal cell carcinoma, roughly one in four, and close to one in five in skin squamous cell carcinoma. Other tumor types had lower rates, suggesting a mutagenesis process—like ultraviolet light—that’s particularly effective at causing these promoter changes in skin.

Were these inherited? That was the next check. Most promoter variants that could be classified—about ninety-three percent—appeared somatic based on their somatic-germline algorithm, and their allele fractions hovered around twenty-five percent, consistent with subclonal acquisition.

To validate this finding in an independent set, they examined The Cancer Genome Atlas melanoma cohort. The same promoter positions appeared there too. Four out of fifty melanomas in that set carried them, around eight percent, and when the analysis was expanded, the overall estimate returned to roughly ten percent with reasonable confidence bounds.

Functionally, several computational predictors pointed to two promoter coordinates—six million forty-eight thousand seven hundred sixty and six million forty-eight thousand eight hundred twenty-four—as the most likely to alter gene regulation. When they investigated whether these promoter mutations tended to co-occur with other known drivers in melanoma, none held up after adjusting for TMB. This indicates a clean, recurrent regulatory hotspot linked to mutation load in skin cancers, acting predominantly on its own.

Stepping back, a few themes emerge. First, a thoughtfully designed panel of about 1.1 megabases can substitute for exome sequencing when the goal is to quantify tumor mutational burden. The evidence is solid: there is a strong correlation against exome in paired tumors, nearly perfect reproducibility on replicates, and simulation studies that clarify why panel size matters—significantly shrinking below half a megabase leads to noisy results.

Second, TMB is a landscape, not a single peak. The median is modest across cancers, but high-TMB clusters emerge in dozens of diseases, including rare ones, and they become more common as patients age. Third, while MSI captures part of the hypermutation story, particularly in gastrointestinal contexts, many high-TMB tumors aren’t MSI-high; other mechanisms—including alterations in mismatch repair genes, POLE mutations, and regulatory hits like the PMS2 promoter mutations—are involved.

There are limits to this approach. As Frampton and colleagues clarify, a targeted panel isn’t the best tool for discovering every neoantigen; by its nature, it only observes a sliver of the genome and won’t capture the entire antigenic catalog. Germline filtering, although stringent, remains an area to refine further, especially for ancestries that have been underrepresented in reference databases.

But for the fundamental, often repeated question in clinics—does this tumor carry enough mutational weight that a checkpoint inhibitor might have a chance?—this approach provides an answer that aligns well with whole-exome truth and can be scaled effectively.

Clinically, this means TMB joins MSI, PD-L1, and specific drivers as one of the key pillars to evaluate. It’s not a kingmaker on its own, but it helps identify patients across a variety of tumor types who might benefit from immunotherapy. It also encourages us to examine genomes with both breadth and nuance: the base-by-base counts, the repair pathways that have failed, and the peculiar regulatory lesion that tips a balance.

As Chalmers and colleagues demonstrated by mapping more than one hundred thousand cancers, when you measure the right factors in the right way, the signal begins to stand out. You obtain a map you can utilize. And in oncology, that map often represents the difference between guessing and making informed choices.

More in Medicine