Language Structure Is Partly Determined by Social Structure

Gary Lupyan, Rick DaleView original
OverviewBalancedalloy voice
Think about the puzzle of human languages for a second. Some are awash in endings and affixes, where a single word can carry tense, case, person, even how you know what you're saying. Others strip all that off and lean on little words and word order to do the job. These languages live in very different social worlds. Some are spoken by millions spread across continents; others by tight-knit communities with few outsiders. The big question is whether those social worlds actually shape the grammar itself. Gary Lupyan and Rick Dale make a bold, testable claim: languages adapt to their social niches. They call it the Linguistic Niche Hypothesis. Picture a continuum. At one end, the exoteric niche: large populations, broad geography, and lots of contact with other languages—think English or Swahili. At the other end, the esoteric niche: small, cohesive groups with fewer outsiders—think Elfdalian or Algonquin. The prediction is straightforward. Exoteric languages, constantly used by strangers and learned by many adults, will shed hard-to-learn inflection and instead express distinctions with separate words. Esoteric languages, passed mainly among children within a community, can sustain—and sometimes favor—rich, redundant morphology. That mechanism matters. Adult second-language learners tend to struggle with opaque inflectional paradigms. If many adults are learning, pressures mount against keeping irregular, highly fused bits of grammar. Over generations, that can tilt a language away from complex inflection and toward clearer, more compositional expressions—particles for negation, separate words for evidentiality, and analytical ways of marking aspect and possession. This is the idea. But can you see it in the data? To find out, Lupyan and Dale stitched together an unusually broad dataset. They started with typological features from the World Atlas of Language Structures—things like how many cases a language marks or how much verbal morphology it packs into a single word. Then they paired this with three nonlinguistic proxies for the social niche: the number of native speakers, the geographic area where the language is used, and a rough measure of contact—the count of other languages whose territories touch or overlap. Population came from Ethnologue, area from digitized language polygons, and neighbors from how those polygons intersect. None of these is perfect. Together, they triangulate exoteric versus esoteric conditions in a way that can be modeled. Crucially, they took the patchwork nature of typological data seriously. No single feature was coded for every language. Across the atlas, a typical feature had data for a couple hundred languages, and coverage varied. So they did two things. First, they ran feature-by-feature analyses using appropriate statistics—logistic regressions for binary features, multinomial models for unordered categories, and standard linear models for continuous measures like inflectional synthesis of the verb. Second, they built a composite morphological complexity score across many features and adjusted it for missingness, dividing by how many features were actually known for each language. Languages with only a handful of entries were excluded. The aim was simple: don't let data gaps masquerade as simplicity or complexity. Now the punchline. Across a sample of two thousand two hundred thirty-six languages, demographic variables predict morphology. Strongly. Population—and to a lesser extent area and number of neighboring languages—significantly predicted twenty-six of twenty-eight morphology-relevant features they examined. When they accounted for shared ancestry by partialing out language family, twenty-three of those relationships held. When they asked whether adding demographics to geography actually helped, the answer was yes: combining the social variables with geographic covariates outperformed geography alone in twenty-two of twenty-eight cases. In plain terms, who speaks a language, where, and around whom tells you something real about how that language builds words. What does that look like on the ground? In bigger, more contact-heavy populations, languages tend to shift information away from bound morphemes and into free words. If a language marks future tense morphologically, it's more likely to be one with fewer speakers; if it uses a separate particle for the future, more speakers. Possession follows a similar pattern, and so does the way languages handle adpositions—whether information is bundled into an affix or sits next to the noun as a separate word. You see the same tilt for negation, evidentiality, modality, and aspect: exoteric ecologies favor lexical strategies. Zooming out, the structural fingerprints line up. Languages in exoteric niches are more isolating rather than fusional, carry fewer case distinctions, show more case syncretism, and strip down agreement across nouns and verbs. They often have less elaborate verb morphology and avoid packing person marking into adpositions. None of this means they're simplistic. It means they move complexity from the realm of paradigms and agreement into transparent, word-by-word composition. You can hear the upside if you've ever learned a language that says "did not go" instead of fusing negation into a tangle of endings. All of this converges in the composite measure. When Lupyan and Dale rolled many features into a single adjusted complexity score, the relationship with population was not just visible, it was statistically overwhelming. Larger speaker populations went hand in hand with lower morphological complexity, and the probability of that pattern being a fluke was vanishingly small—on the order of five in one hundred thousand. The feature-by-feature analyses and the one-number summary pointed to the same place. But are we just tracing family trees and map coordinates? The authors worked hard to rule that out. They included latitude and longitude as covariates, and in some analyses, treated continent as a random effect, so regional clustering wouldn't drive results. They also leaned into Galton's problem—the fact that languages are related and share history—by doing two more checks. First, they aggregated within families and still found a robust link. For a core measure of verbal inflectional synthesis, the within-family association with population was sky-high when they averaged by geographic location—think a Pearson correlation on the order of point nine. Across families, the link was moderate but clear as well; when you average by family for that same measure, the correlation sits around one-half. Second, they performed a Monte Carlo scramble: randomize which demographic profiles go with which languages, but only within families, and see what happens. What happened is that the predictive power of population dropped for most features—twenty-two out of the twenty-eight. That's a strong sign the observed pattern isn't just an echo of shared ancestry. There's also the reality check of outliers. Lingua francas—the languages that spread widely and stitch together diverse groups—often sit even further below the within-family trend than you'd expect. That's what exoteric pressure looks like in the wild: extra simplification beyond what related languages show. And yes, there are exceptions. Some families with small ranges, like the Australian family, don't follow the population effect as cleanly. Those exceptions are important—they tell you this isn't destiny; it's pressure. Rival explanations get a fair hearing. Maybe tiny communities just drift into complicated paradigms because change is buffered within small groups. Maybe languages with few speakers are more efficient if they pack lots of information into single words. Or maybe when a language spreads through formal schooling, prescriptive norms sand down irregularities. Each of these can explain a corner of the map. What they don't do, when you put them under the same modeling and controls, is produce the broad, multi-feature pattern you see across thousands of languages and continents. The demographic predictors keep showing up, and they keep showing up after you've accounted for history and place. So what's the engine under the hood? Learnability. Lupyan and Dale sketch a simple model of how the proportion of adult second-language learners changes the fitness of different grammatical systems. Imagine a curve where fitness reflects how easily a community can acquire and use a language. As the share of adult learners rises, the peak of that curve moves toward grammars that minimize the number of bound distinctions and reduce redundancy. Fewer inflectional categories, cleaner one-meaning-per-word pairings. In communities where almost everyone learns as a child, the peak shifts the other way. Richer, more redundant morphology isn't a bug; it's a feature that helps kids generalize from less input. That is, redundancy can support robustness in acquisition when first language learning dominates. They back that mechanism with an independent check. Using a large translation-based dataset spanning just over a hundred languages, they tested the link between speaker population and grammatical redundancy. The association was highly significant, with a tiny p-value well below one in one thousand. That external line of evidence matches the main finding: languages with larger speaker populations tend to be less morphologically specified and less redundant in their bound morphology, relying more on lexical cues instead. It's worth pausing on the texture of those demographic predictors themselves. Population, area, and number of neighboring languages aren't independent—each correlates with the others by about half. But including all three in the models matters. It keeps any single proxy from pulling the result, and together they capture different facets of the exoteric—esoteric spectrum: how many people need to learn the system, how far it stretches, and how messy the contact zone is. All three contribute, with population consistently the strongest. What does this mean for how we think about grammar? It suggests that morphological complexity is not a fixed property—a language's soul—so much as a moving balance between learnability pressures and communicative needs shaped by a community's social ecology. Exoteric languages trend toward lexical transparency. Esoteric languages sustain morphological richness and redundancy. Both are adaptive. As Lupyan and Dale show, you can predict aspects of grammar from simple facts about who's speaking, where, and in what kind of contact network. It also nudges typology toward a more ecological framing. If exoteric pressure pushes meaning into separate words and increases compositional transparency, we should expect certain co-occurring patterns: fewer agreement morphemes, fewer case endings, and more particles handling tense, negation, and evidentiality. That's exactly what emerges in the data. It means population-based variables aren't just background—they have real predictive bite for grammatical profiles. Of course, there are limits. The World Atlas of Language Structures is a marvel, but it's sparse by design, and not every feature of interest is coded for every language. Some links are sensitive to how you slice the sample. While the learnability story is compelling, the field still needs direct experimental evidence that specific bits of morphology are harder for adults than for children in the ways this theory assumes. Lupyan and Dale say as much. The hypothesis is a strong first pass, not the last word. Still, the core takeaway holds. Linguistic structure is partly determined by social structure. Run the tape forward: as languages spread, as adult learners come in, inflectional paradigms simplify and lexical strategies take on more of the load. In denser, more insular communities, morphology stays elaborate and redundant, suited to a world where kids are the primary learners and shared context does more of the work. It's a neat synthesis—social worlds leaving fingerprints on grammar, and grammar, in turn, adapting to fit the niche. This lecture was created by ennepō. Go to ennepo dot A I to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Think about the puzzle of human languages for a second. Some are awash in endings and affixes, where a single word can carry tense, case, person, even how you know what you're saying. Others strip all that off and lean on little words and word order to do the job.

These languages live in very different social worlds. Some are spoken by millions spread across continents; others by tight-knit communities with few outsiders. The big question is whether those social worlds actually shape the grammar itself.

Gary Lupyan and Rick Dale make a bold, testable claim: languages adapt to their social niches. They call it the Linguistic Niche Hypothesis. Picture a continuum.

At one end, the exoteric niche: large populations, broad geography, and lots of contact with other languages—think English or Swahili. At the other end, the esoteric niche: small, cohesive groups with fewer outsiders—think Elfdalian or Algonquin. The prediction is straightforward.

Exoteric languages, constantly used by strangers and learned by many adults, will shed hard-to-learn inflection and instead express distinctions with separate words. Esoteric languages, passed mainly among children within a community, can sustain—and sometimes favor—rich, redundant morphology.

That mechanism matters. Adult second-language learners tend to struggle with opaque inflectional paradigms. If many adults are learning, pressures mount against keeping irregular, highly fused bits of grammar.

Over generations, that can tilt a language away from complex inflection and toward clearer, more compositional expressions—particles for negation, separate words for evidentiality, and analytical ways of marking aspect and possession. This is the idea. But can you see it in the data?

To find out, Lupyan and Dale stitched together an unusually broad dataset. They started with typological features from the World Atlas of Language Structures—things like how many cases a language marks or how much verbal morphology it packs into a single word. Then they paired this with three nonlinguistic proxies for the social niche: the number of native speakers, the geographic area where the language is used, and a rough measure of contact—the count of other languages whose territories touch or overlap.

Population came from Ethnologue, area from digitized language polygons, and neighbors from how those polygons intersect. None of these is perfect. Together, they triangulate exoteric versus esoteric conditions in a way that can be modeled.

Crucially, they took the patchwork nature of typological data seriously. No single feature was coded for every language. Across the atlas, a typical feature had data for a couple hundred languages, and coverage varied.

So they did two things. First, they ran feature-by-feature analyses using appropriate statistics—logistic regressions for binary features, multinomial models for unordered categories, and standard linear models for continuous measures like inflectional synthesis of the verb. Second, they built a composite morphological complexity score across many features and adjusted it for missingness, dividing by how many features were actually known for each language.

Languages with only a handful of entries were excluded. The aim was simple: don't let data gaps masquerade as simplicity or complexity.

Now the punchline. Across a sample of two thousand two hundred thirty-six languages, demographic variables predict morphology. Strongly.

Population—and to a lesser extent area and number of neighboring languages—significantly predicted twenty-six of twenty-eight morphology-relevant features they examined. When they accounted for shared ancestry by partialing out language family, twenty-three of those relationships held. When they asked whether adding demographics to geography actually helped, the answer was yes: combining the social variables with geographic covariates outperformed geography alone in twenty-two of twenty-eight cases.

In plain terms, who speaks a language, where, and around whom tells you something real about how that language builds words.

What does that look like on the ground? In bigger, more contact-heavy populations, languages tend to shift information away from bound morphemes and into free words. If a language marks future tense morphologically, it's more likely to be one with fewer speakers; if it uses a separate particle for the future, more speakers.

Possession follows a similar pattern, and so does the way languages handle adpositions—whether information is bundled into an affix or sits next to the noun as a separate word. You see the same tilt for negation, evidentiality, modality, and aspect: exoteric ecologies favor lexical strategies.

Zooming out, the structural fingerprints line up. Languages in exoteric niches are more isolating rather than fusional, carry fewer case distinctions, show more case syncretism, and strip down agreement across nouns and verbs. They often have less elaborate verb morphology and avoid packing person marking into adpositions.

None of this means they're simplistic. It means they move complexity from the realm of paradigms and agreement into transparent, word-by-word composition. You can hear the upside if you've ever learned a language that says "did not go" instead of fusing negation into a tangle of endings.

All of this converges in the composite measure. When Lupyan and Dale rolled many features into a single adjusted complexity score, the relationship with population was not just visible, it was statistically overwhelming. Larger speaker populations went hand in hand with lower morphological complexity, and the probability of that pattern being a fluke was vanishingly small—on the order of five in one hundred thousand.

The feature-by-feature analyses and the one-number summary pointed to the same place.

But are we just tracing family trees and map coordinates? The authors worked hard to rule that out. They included latitude and longitude as covariates, and in some analyses, treated continent as a random effect, so regional clustering wouldn't drive results.

They also leaned into Galton's problem—the fact that languages are related and share history—by doing two more checks. First, they aggregated within families and still found a robust link. For a core measure of verbal inflectional synthesis, the within-family association with population was sky-high when they averaged by geographic location—think a Pearson correlation on the order of point nine.

Across families, the link was moderate but clear as well; when you average by family for that same measure, the correlation sits around one-half. Second, they performed a Monte Carlo scramble: randomize which demographic profiles go with which languages, but only within families, and see what happens. What happened is that the predictive power of population dropped for most features—twenty-two out of the twenty-eight.

That's a strong sign the observed pattern isn't just an echo of shared ancestry.

There's also the reality check of outliers. Lingua francas—the languages that spread widely and stitch together diverse groups—often sit even further below the within-family trend than you'd expect. That's what exoteric pressure looks like in the wild: extra simplification beyond what related languages show.

And yes, there are exceptions. Some families with small ranges, like the Australian family, don't follow the population effect as cleanly. Those exceptions are important—they tell you this isn't destiny; it's pressure.

Rival explanations get a fair hearing. Maybe tiny communities just drift into complicated paradigms because change is buffered within small groups. Maybe languages with few speakers are more efficient if they pack lots of information into single words.

Or maybe when a language spreads through formal schooling, prescriptive norms sand down irregularities. Each of these can explain a corner of the map. What they don't do, when you put them under the same modeling and controls, is produce the broad, multi-feature pattern you see across thousands of languages and continents.

The demographic predictors keep showing up, and they keep showing up after you've accounted for history and place.

So what's the engine under the hood? Learnability. Lupyan and Dale sketch a simple model of how the proportion of adult second-language learners changes the fitness of different grammatical systems.

Imagine a curve where fitness reflects how easily a community can acquire and use a language. As the share of adult learners rises, the peak of that curve moves toward grammars that minimize the number of bound distinctions and reduce redundancy. Fewer inflectional categories, cleaner one-meaning-per-word pairings.

In communities where almost everyone learns as a child, the peak shifts the other way. Richer, more redundant morphology isn't a bug; it's a feature that helps kids generalize from less input. That is, redundancy can support robustness in acquisition when first language learning dominates.

They back that mechanism with an independent check. Using a large translation-based dataset spanning just over a hundred languages, they tested the link between speaker population and grammatical redundancy. The association was highly significant, with a tiny p-value well below one in one thousand.

That external line of evidence matches the main finding: languages with larger speaker populations tend to be less morphologically specified and less redundant in their bound morphology, relying more on lexical cues instead.

It's worth pausing on the texture of those demographic predictors themselves. Population, area, and number of neighboring languages aren't independent—each correlates with the others by about half. But including all three in the models matters.

It keeps any single proxy from pulling the result, and together they capture different facets of the exoteric—esoteric spectrum: how many people need to learn the system, how far it stretches, and how messy the contact zone is. All three contribute, with population consistently the strongest.

What does this mean for how we think about grammar? It suggests that morphological complexity is not a fixed property—a language's soul—so much as a moving balance between learnability pressures and communicative needs shaped by a community's social ecology. Exoteric languages trend toward lexical transparency.

Esoteric languages sustain morphological richness and redundancy. Both are adaptive. As Lupyan and Dale show, you can predict aspects of grammar from simple facts about who's speaking, where, and in what kind of contact network.

It also nudges typology toward a more ecological framing. If exoteric pressure pushes meaning into separate words and increases compositional transparency, we should expect certain co-occurring patterns: fewer agreement morphemes, fewer case endings, and more particles handling tense, negation, and evidentiality. That's exactly what emerges in the data.

It means population-based variables aren't just background—they have real predictive bite for grammatical profiles.

Of course, there are limits. The World Atlas of Language Structures is a marvel, but it's sparse by design, and not every feature of interest is coded for every language. Some links are sensitive to how you slice the sample.

While the learnability story is compelling, the field still needs direct experimental evidence that specific bits of morphology are harder for adults than for children in the ways this theory assumes. Lupyan and Dale say as much. The hypothesis is a strong first pass, not the last word.

Still, the core takeaway holds. Linguistic structure is partly determined by social structure. Run the tape forward: as languages spread, as adult learners come in, inflectional paradigms simplify and lexical strategies take on more of the load.

In denser, more insular communities, morphology stays elaborate and redundant, suited to a world where kids are the primary learners and shared context does more of the work. It's a neat synthesis—social worlds leaving fingerprints on grammar, and grammar, in turn, adapting to fit the niche.

This lecture was created by ennepō.

Go to ennepo dot A I to Discover, Create and Follow the latest research in your field.

Read when you can. Listen when you want to.

More in Social Sciences