Phenotypic and molecular characterization of the claudin-low intrinsic subtype of breast cancer

Aleix Prat, Joel S. Parker, Olga Karginova, Cheng Fan, Chad Livasy, Jason I. Herschkowitz, Xiaping He, Charles M. PerouView original
OverviewBalancedalloy voice
Picture the map of breast cancer most of us learned: luminal A, luminal B, HER2-enriched, and basal-like. Prat and colleagues argued there’s a missing neighborhood on that map, a place that looks different under the microscope, behaves differently in patients, and carries a distinct molecular signature. They called it claudin-low. It sits near basal-like in the clustering tree, but when you walk the streets, it’s a different city: fewer cell-cell junctions, more mesenchymal traits, and a strong whiff of stemness. Clinically, claudin-low lives largely inside the triple-negative arena, but not all the way inside basal-like. Across several cohorts—UNC337, NKI295, and the MD Anderson neoadjuvant set—it makes up about 7 to 14 percent of all breast cancers. Within triple-negative disease, most tumors are basal-like, roughly 39 to 54 percent, yet a substantial slice, about 25 to 39 percent, are claudin-low. That’s not a sliver. It’s enough to reshape how we think about triple-negative breast cancer as a single bucket. These tumors also cycle more slowly than basal-like, with Ki67 significantly lower, so they can look deceptively quiet even as they portend trouble. What sets claudin-low apart at the molecular level is the collapse of epithelial adhesion. Think of tight junctions like Velcro strips between cells; claudin-low tumors tear them off. E-cadherin drops. Claudin 3, 4, and 7 drop. In UNC337, immunohistochemistry showed low-to-absent E-cadherin in 45 percent of claudin-low tumors compared to 11 percent of basal-like, and claudin-3 low-to-absent in 59 versus 11 percent. Compared to all other subtypes combined, those same losses—45 and 59 percent—still stand out. The flip side rises: immune and stromal programs are high, so you see a tumor bed thick with signals from fibroblasts, lymphocytes, and extracellular matrix. This is classic epithelial-to-mesenchymal transition, or EMT, but sustained and system-wide. The transcription factors that drive EMT—SNAI1 and SNAI2, TWIST1 and TWIST2, ZEB1 and ZEB2—light up. Vimentin, a cytoskeletal marker of mesenchyme, comes along for the ride, and hypoxia pathways tag in. Compared with basal-like, the claudin-low group shows a broad canvas of differences: more than a thousand genes up, a few hundred down, with immune signaling, extracellular matrix, and cell migration showing up repeatedly. It’s not a tweak. It’s a shift in cell identity. That shift carries a functional consequence: the hallmarks of tumor-initiating cells. Across subtypes, claudin-low is where you find the highest and most consistent enrichment for stem-like signatures. The CD44-high, CD24-low profile? It’s prominent. The CD49f-positive, EpCAM-low profile? The same. ALDH1A1, a stem-associated enzyme, trends up as well. Three different breast stem cell-like gene sets, derived from distinct studies with little gene overlap, each pour into claudin-low with striking enrichment. When three routes lead to the same hilltop, you look up and pay attention. Now, how did they call this subtype with confidence across datasets, cell lines, and even mouse tumors? Here’s the clever bit. Prat’s team built a claudin-low centroid predictor from nine unmistakably claudin-low cell lines. In plain language, they took the average expression fingerprint of those nine lines and asked, for any new sample, "Are you closer to this fingerprint or to everyone else?" When they applied it to the UNC337 tumors, 37 of 337—about 11 percent—landed in claudin-low. Measured against a clustering-based gold standard, the predictor’s sensitivity was 87.5 percent and its specificity 97.0 percent. Some samples that PAM50 had labeled differently, including several basal-like, moved into claudin-low with this cell-line anchor. A reminder that labels are only as good as the lens you use. They pushed that lens across species. In a bank of genetically engineered mouse models, every mouse tumor that the predictor called claudin-low fell into a mesenchymal-heavy category dubbed Group II. Those mouse tumors carried EMT and stem-like signatures that lined up with the human claudin-low pattern. Normal mouse mammary tissues didn’t get misclassified as claudin-low, which helps separate tumor-intrinsic biology from noise. For a subtype grounded in loss of epithelial features, that guardrail matters. To place claudin-low on a developmental timeline, the team borrowed an axis that runs from mammary stem cells to luminal progenitors to mature luminal cells, as described by Lim and colleagues. They oriented tumors along that axis using distance-weighted discrimination—imagine projecting each tumor’s expression profile onto a line between stem-like and mature luminal reference points. Low scores equate to less differentiation. When they did this, claudin-low clustered near the undifferentiated end. And those low differentiation scores weren’t just labels—they tracked with outcomes. In UNC337, a low score associated with a hazard ratio of 2.83 for relapse-free survival and 5.66 for overall survival. In the NKI295 cohort, the separation was even starker, with hazard ratios of 4.71 and 17.98. Those are big gaps. They translate the biology into prognosis. What about the laboratory workhorses we use to model cancer? Neve’s canonical panel of breast cancer cell lines includes a claudin-low set—names you may know: MDA-MB-231, SUM159PT, Hs578T, and BT549. Across 52 lines, nine wore the claudin-low badge in this framework. When cells were pushed into mammosphere culture, which enriches for stem-like states, the fraction of claudin-low cases climbed; in one set of 14 mammospheres, six were claudin-low. Even the broader NCI-60 panel held four claudin-low lines. Consistency across testbeds gives you confidence the subtype isn’t a data artifact. Protein-level and single-cell-like readouts tell the same story. In a cohort of tumors with dual staining for epithelial keratins 5 and 19 and for vimentin, a third—28 of 86—showed both marks at once. Almost nine in ten of those dual-positive tumors were claudin-low or basal-like. Zoom in on claudin-low and the dual identity becomes even more common: around 55 percent show both epithelial and mesenchymal markers by immunofluorescence, compared with about 26 percent across other subtypes. Flow sorting puts names to faces. In the SUM149PT line, the CD49f-positive, EpCAM-low compartment carries a mesenchymal, claudin-low-like transcriptional profile—high interleukin-6, CXCL1, vascular endothelial growth factor A, vimentin, and SNAI1; low E-cadherin, claudin-7, keratin-19, and CD24. And a small slice of those cells, about 5 to 10 percent, will differentiate toward a more basal-like state. Plasticity is part of the picture. So how do these tumors behave in patients? They sit in a tricky middle. On long-term outcomes, claudin-low does worse than luminal A and aligns with the other poor-prognosis subtypes—basal-like, HER2-enriched, and luminal B. In the neoadjuvant setting at MD Anderson, claudin-low tumors had a pathologic complete response rate of 38.9 percent to anthracycline and taxane chemotherapy. Basal-like tumors hit 73.3 percent. Luminal A and B lag far behind both. Here’s a wrinkle: because claudin-low can be miscalled as basal-like by some classifiers, mixing the two can inflate the apparent response of basal-like disease. When separated cleanly, the difference is obvious—basal-like retains the highest chemo-sensitivity, claudin-low shows intermediate response, and yet both share poor survival. That’s the clinical paradox the biology explains. If you peek back at the microscope, the histology fits. Claudin-low tumors often show metaplastic or medullary features—less gland-forming, more spindled or squamous differentiation—echoing the EMT and stromal signatures in the gene expression. They’re not monolithic; there’s diversity within that envelope. But it’s a different envelope from the classic luminal or basal-like glands. There’s a caveat worth underlining. When a tumor’s gene expression looks stromal, you have to ask: are you measuring the tumor or the neighborhood? Prat’s team anticipated that. They trained their claudin-low predictor on cell lines, which lets them focus on the epithelial compartment, and then validated it across tumors. The sensitivity and specificity held strong against a clustering-based gold standard. They still caution, appropriately, that tumor cellularity and careful separation of tumor cells from fibroblastic stroma are critical to avoid misclassification. Put simply: don’t call the town hall a factory just because you sampled the industrial park next door. Across systems—human tumors, cultured lines, and genetically engineered mouse models—the same silhouette appears. Loss of claudins and E-cadherin. Rise of EMT drivers and vimentin. Immune and stromal signals bleeding into the profile. A tilt toward tumor-initiating cell states that places these tumors at the undifferentiated end of the mammary hierarchy described by Lim and colleagues. Even normal human mammary subpopulations sorted by Raouf’s group behaved as you’d expect in this framework: the bipotent progenitors mapped claudin-low by the predictor with perfect concordance in a small test set, eight out of eight. That’s the kind of cross-check that keeps a subtype from collapsing under its own complexity. So what changes if you recognize claudin-low as its own place on the map? First, you stop assuming all triple-negative tumors are basal-like. That helps with prognosis and with expectations around chemotherapy response. Second, you get a clearer link between an undifferentiated, EMT-rich state and clinical outcomes—those striking hazard ratios aren’t abstractions; they’re patients. And third, you gain aligned models—specific cell lines and mouse tumors—that actually look like the disease you’re trying to study, which is rarer than we admit in cancer biology. Where does this go next? Carefully, and with focus. The claudin-low predictor gives a way to enrich trials for this biology, test whether EMT-targeted or stemness-targeted strategies move the needle, and explore immunomodulatory approaches tuned to the heavy immune and stromal context these tumors inhabit. It also invites deeper dissection of plasticity—the fact that a fraction of claudin-low-like cells can drift toward basal-like—and whether intercepting that drift matters. But the big message is already here. As Prat and colleagues showed, claudin-low isn’t a blur between subtypes. It’s a distinct destination—poorly differentiated, EMT-heavy and stem cell-heavy, clinically tough—that we can now name, measure, and model. And once you can point to a place on the map, you can start building roads in and out. That’s how precision medicine actually gets precise.

Picture the map of breast cancer most of us learned: luminal A, luminal B, HER2-enriched, and basal-like. Prat and colleagues argued there’s a missing neighborhood on that map, a place that looks different under the microscope, behaves differently in patients, and carries a distinct molecular signature. They called it claudin-low.

It sits near basal-like in the clustering tree, but when you walk the streets, it’s a different city: fewer cell-cell junctions, more mesenchymal traits, and a strong whiff of stemness.

Clinically, claudin-low lives largely inside the triple-negative arena, but not all the way inside basal-like. Across several cohorts—UNC337, NKI295, and the MD Anderson neoadjuvant set—it makes up about 7 to 14 percent of all breast cancers. Within triple-negative disease, most tumors are basal-like, roughly 39 to 54 percent, yet a substantial slice, about 25 to 39 percent, are claudin-low.

That’s not a sliver. It’s enough to reshape how we think about triple-negative breast cancer as a single bucket. These tumors also cycle more slowly than basal-like, with Ki67 significantly lower, so they can look deceptively quiet even as they portend trouble.

What sets claudin-low apart at the molecular level is the collapse of epithelial adhesion. Think of tight junctions like Velcro strips between cells; claudin-low tumors tear them off. E-cadherin drops.

Claudin 3, 4, and 7 drop. In UNC337, immunohistochemistry showed low-to-absent E-cadherin in 45 percent of claudin-low tumors compared to 11 percent of basal-like, and claudin-3 low-to-absent in 59 versus 11 percent. Compared to all other subtypes combined, those same losses—45 and 59 percent—still stand out.

The flip side rises: immune and stromal programs are high, so you see a tumor bed thick with signals from fibroblasts, lymphocytes, and extracellular matrix.

This is classic epithelial-to-mesenchymal transition, or EMT, but sustained and system-wide. The transcription factors that drive EMT—SNAI1 and SNAI2, TWIST1 and TWIST2, ZEB1 and ZEB2—light up. Vimentin, a cytoskeletal marker of mesenchyme, comes along for the ride, and hypoxia pathways tag in.

Compared with basal-like, the claudin-low group shows a broad canvas of differences: more than a thousand genes up, a few hundred down, with immune signaling, extracellular matrix, and cell migration showing up repeatedly. It’s not a tweak. It’s a shift in cell identity.

That shift carries a functional consequence: the hallmarks of tumor-initiating cells. Across subtypes, claudin-low is where you find the highest and most consistent enrichment for stem-like signatures. The CD44-high, CD24-low profile?

It’s prominent. The CD49f-positive, EpCAM-low profile? The same.

ALDH1A1, a stem-associated enzyme, trends up as well. Three different breast stem cell-like gene sets, derived from distinct studies with little gene overlap, each pour into claudin-low with striking enrichment. When three routes lead to the same hilltop, you look up and pay attention.

Now, how did they call this subtype with confidence across datasets, cell lines, and even mouse tumors? Here’s the clever bit. Prat’s team built a claudin-low centroid predictor from nine unmistakably claudin-low cell lines.

In plain language, they took the average expression fingerprint of those nine lines and asked, for any new sample, "Are you closer to this fingerprint or to everyone else?" When they applied it to the UNC337 tumors, 37 of 337—about 11 percent—landed in claudin-low. Measured against a clustering-based gold standard, the predictor’s sensitivity was 87.5 percent and its specificity 97.0 percent. Some samples that PAM50 had labeled differently, including several basal-like, moved into claudin-low with this cell-line anchor. A reminder that labels are only as good as the lens you use.

They pushed that lens across species. In a bank of genetically engineered mouse models, every mouse tumor that the predictor called claudin-low fell into a mesenchymal-heavy category dubbed Group II. Those mouse tumors carried EMT and stem-like signatures that lined up with the human claudin-low pattern.

Normal mouse mammary tissues didn’t get misclassified as claudin-low, which helps separate tumor-intrinsic biology from noise. For a subtype grounded in loss of epithelial features, that guardrail matters.

To place claudin-low on a developmental timeline, the team borrowed an axis that runs from mammary stem cells to luminal progenitors to mature luminal cells, as described by Lim and colleagues. They oriented tumors along that axis using distance-weighted discrimination—imagine projecting each tumor’s expression profile onto a line between stem-like and mature luminal reference points. Low scores equate to less differentiation.

When they did this, claudin-low clustered near the undifferentiated end. And those low differentiation scores weren’t just labels—they tracked with outcomes. In UNC337, a low score associated with a hazard ratio of 2.83 for relapse-free survival and 5.66 for overall survival.

In the NKI295 cohort, the separation was even starker, with hazard ratios of 4.71 and 17.98. Those are big gaps. They translate the biology into prognosis.

What about the laboratory workhorses we use to model cancer? Neve’s canonical panel of breast cancer cell lines includes a claudin-low set—names you may know: MDA-MB-231, SUM159PT, Hs578T, and BT549. Across 52 lines, nine wore the claudin-low badge in this framework.

When cells were pushed into mammosphere culture, which enriches for stem-like states, the fraction of claudin-low cases climbed; in one set of 14 mammospheres, six were claudin-low. Even the broader NCI-60 panel held four claudin-low lines. Consistency across testbeds gives you confidence the subtype isn’t a data artifact.

Protein-level and single-cell-like readouts tell the same story. In a cohort of tumors with dual staining for epithelial keratins 5 and 19 and for vimentin, a third—28 of 86—showed both marks at once. Almost nine in ten of those dual-positive tumors were claudin-low or basal-like.

Zoom in on claudin-low and the dual identity becomes even more common: around 55 percent show both epithelial and mesenchymal markers by immunofluorescence, compared with about 26 percent across other subtypes. Flow sorting puts names to faces. In the SUM149PT line, the CD49f-positive, EpCAM-low compartment carries a mesenchymal, claudin-low-like transcriptional profile—high interleukin-6, CXCL1, vascular endothelial growth factor A, vimentin, and SNAI1; low E-cadherin, claudin-7, keratin-19, and CD24.

And a small slice of those cells, about 5 to 10 percent, will differentiate toward a more basal-like state. Plasticity is part of the picture.

So how do these tumors behave in patients? They sit in a tricky middle. On long-term outcomes, claudin-low does worse than luminal A and aligns with the other poor-prognosis subtypes—basal-like, HER2-enriched, and luminal B.

In the neoadjuvant setting at MD Anderson, claudin-low tumors had a pathologic complete response rate of 38.9 percent to anthracycline and taxane chemotherapy. Basal-like tumors hit 73.3 percent. Luminal A and B lag far behind both.

Here’s a wrinkle: because claudin-low can be miscalled as basal-like by some classifiers, mixing the two can inflate the apparent response of basal-like disease. When separated cleanly, the difference is obvious—basal-like retains the highest chemo-sensitivity, claudin-low shows intermediate response, and yet both share poor survival. That’s the clinical paradox the biology explains.

If you peek back at the microscope, the histology fits. Claudin-low tumors often show metaplastic or medullary features—less gland-forming, more spindled or squamous differentiation—echoing the EMT and stromal signatures in the gene expression. They’re not monolithic; there’s diversity within that envelope. But it’s a different envelope from the classic luminal or basal-like glands.

There’s a caveat worth underlining. When a tumor’s gene expression looks stromal, you have to ask: are you measuring the tumor or the neighborhood? Prat’s team anticipated that.

They trained their claudin-low predictor on cell lines, which lets them focus on the epithelial compartment, and then validated it across tumors. The sensitivity and specificity held strong against a clustering-based gold standard. They still caution, appropriately, that tumor cellularity and careful separation of tumor cells from fibroblastic stroma are critical to avoid misclassification.

Put simply: don’t call the town hall a factory just because you sampled the industrial park next door.

Across systems—human tumors, cultured lines, and genetically engineered mouse models—the same silhouette appears. Loss of claudins and E-cadherin. Rise of EMT drivers and vimentin.

Immune and stromal signals bleeding into the profile. A tilt toward tumor-initiating cell states that places these tumors at the undifferentiated end of the mammary hierarchy described by Lim and colleagues. Even normal human mammary subpopulations sorted by Raouf’s group behaved as you’d expect in this framework: the bipotent progenitors mapped claudin-low by the predictor with perfect concordance in a small test set, eight out of eight.

That’s the kind of cross-check that keeps a subtype from collapsing under its own complexity.

So what changes if you recognize claudin-low as its own place on the map? First, you stop assuming all triple-negative tumors are basal-like. That helps with prognosis and with expectations around chemotherapy response.

Second, you get a clearer link between an undifferentiated, EMT-rich state and clinical outcomes—those striking hazard ratios aren’t abstractions; they’re patients. And third, you gain aligned models—specific cell lines and mouse tumors—that actually look like the disease you’re trying to study, which is rarer than we admit in cancer biology.

Where does this go next? Carefully, and with focus. The claudin-low predictor gives a way to enrich trials for this biology, test whether EMT-targeted or stemness-targeted strategies move the needle, and explore immunomodulatory approaches tuned to the heavy immune and stromal context these tumors inhabit.

It also invites deeper dissection of plasticity—the fact that a fraction of claudin-low-like cells can drift toward basal-like—and whether intercepting that drift matters.

But the big message is already here. As Prat and colleagues showed, claudin-low isn’t a blur between subtypes. It’s a distinct destination—poorly differentiated, EMT-heavy and stem cell-heavy, clinically tough—that we can now name, measure, and model.

And once you can point to a place on the map, you can start building roads in and out. That’s how precision medicine actually gets precise.

More in Neuroscience