Personality, Gender, and Age in the Language of Social MediaThe Open-Vocabulary Approach

H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Lukasz Dziurzynski, Stephanie M. Ramones, Megha Agrawal, Achal Shah, Michał Kosiński, David Stillwell, Martin E. P. Seligman, Lyle UngarView original
OverviewBalancedhelen voice
Seven hundred million words are sitting in a database. This is not a century of literature or a curated archive — just Facebook status updates from seventy-five thousand people, written over a few years without any thought of being studied. Schwartz and colleagues at the University of Pennsylvania let that data speak first. They imposed no categories, asked no leading questions, and what came back was a detailed portrait of who we are — our age, our gender, our deepest personality traits — written invisibly into the words we choose without ever knowing we were being read. To understand why this matters, you have to know what researchers had been doing before. The standard approach to language and personality is what Schwartz and colleagues call closed vocabulary. You start by asking humans to decide which words matter, sorting thousands of terms into fixed categories, and then counting how often those categories appear in someone's text. The best-known tool for this is Linguistic Inquiry and Word Count, or LIWC, a manually constructed lexicon with sixty-four categories covering everything from pronoun use to emotion words to family references. It's a serious, carefully built instrument. But it has a fundamental limitation: if you only look for the concepts you already put in the dictionary, you miss everything else. The team cites a concrete example. The word "hug" correlates with agreeableness — but LIWC has no physical-affection category, so it simply vanishes from the analysis. Because the categories were built by human judges working from intuition, emergent patterns the data might reveal, such as a cluster of words tied to partying that distinguishes extraverts, won’t appear if nobody thought to include them. The numbers make this concrete. LIWC-based models predicted gender at seventy-eight point four percent accuracy and correlated with age at a correlation coefficient of zero point sixty-five. For personality, LIWC correlations clustered around zero point twenty-one to zero point twenty-nine across the Big Five traits. Those are not bad numbers. But the open-vocabulary approach beats all of them, sometimes substantially. The method Schwartz and colleagues call differential language analysis, or DLA, works in three steps. First, extract every word, phrase, and automatically generated topic from the text directly — no human curation. From fifteen point four million Facebook messages, that produced twenty-four thousand five hundred thirty word and phrase features, filtered to keep only those used by at least one percent of subjects. On top of that, they ran Latent Dirichlet Allocation, or LDA, a topic-modeling technique that groups words into themes that the data itself reveals — not themes a researcher decided to look for. Second, test each feature for association with the target outcome using regression, controlling for covariates like age and gender so each result reflects a clean relationship. Because tens of thousands of features are being tested simultaneously, they apply strict Bonferroni correction, setting the significance threshold at a p-value below zero point zero zero one. Third, visualize the results with differential word clouds, where the size of each word reflects the strength of its correlation with the trait — not just how frequently it appears overall. The results for gender and age are the clearest demonstration of what the method can do. The sample had forty-six thousand four hundred twelve females and twenty-eight thousand two hundred forty-seven males. Using language features alone, the best predictive model reached ninety-one point nine percent out-of-sample accuracy. The previous best for language-only gender prediction was eighty-eight percent on Twitter. That jump comes from specific, readable signals. Females used more emotion words, more first-person singular pronouns, and more social expressions like "love you" and heart emoticons. Males used more swear words and object references — "xbox" shows up as a significant predictor. One particularly striking micro-pattern: males were far more likely to precede a partner's title with the possessive "my." The phrases "my wife" and "my girlfriend" are strong male predictors. Females used "husband" and "boyfriend" without the possessive, or preceded them with other words — "amazing," "her" — so "my husband" simply didn't emerge as a parallel female signal. Age tells a story about life stages, not just vocabulary. The paper divided users into four groups — thirteen to eighteen, nineteen to twenty-two, twenty-three to twenty-nine, and thirty and older — and the linguistic shift across those groups tracks a recognizable arc: school to college to work to family. The youngest users wrote in emoticons and internet shorthand; the college cohort produced language around "semester," "drunk," and "hangover"; the twenty-something working group wrote about "at work" and "new job"; older users wrote about family. Across the full age range, the word "we" increases approximately linearly after age twenty-two while "I" monotonically decreases — a language-indexed shift toward social integration across the lifespan. Personality is where the findings get genuinely revealing. The Big Five — extraversion, agreeableness, conscientiousness, neuroticism, and openness — are the standard framework in personality psychology, each representing a dimension along which people reliably differ. For each trait, DLA found specific, interpretable linguistic fingerprints. Neuroticism is the starkest. People who score high on neuroticism — emotional instability, anxiety, tendency toward negative affect — disproportionately use the phrase "sick of" and the word "depressed." Their language skews toward negative emotion and first-person singular pronouns, a pattern consistent with the self-focused, ruminative thinking that characterizes the trait. Openness to experience shows up in intellectual vocabulary, references to arts and culture, and a curiosity-inflected range of words that reflects engagement with ideas. Extraversion maps onto social language, positive emotions, and, notably, exclamation marks. Introverts, by contrast, show interest in topics like anime, Pokémon, and manga, along with Japanese emoticons. That kind of specificity is exactly what no closed lexicon would have predicted. Conscientiousness brings words tied to schedules, achievement, and self-discipline — language that reflects organized, goal-directed behavior. Agreeableness maps onto warm, relational language, positivity directed outward rather than inward. One finding the team specifically flags as hypothesis-generating: words tied to physical activity and social participation — "basketball," "snowboarding," "church," "meetings" — correlate with emotional stability, the low end of neuroticism. The implication, carefully framed as a hypothesis rather than a conclusion, is that an active, socially embedded life may be linguistically and psychologically associated with stability. That is the kind of connection a theory-first approach wouldn't have surfaced because nobody thought to code a "physically active lifestyle" category into LIWC. The open-vocabulary model's predictive accuracy for personality consistently outperformed LIWC. LIWC correlations across the Big Five ranged from about zero point twenty-one to zero point twenty-nine; the best open-vocabulary model reached zero point forty-two for extraversion. For gender, the jump from seventy-eight point four percent to ninety-one point nine percent is the clearest summary of what the method adds. Some of the most interesting findings are the ones nobody would have thought to look for at all. People living in high elevations mention mountains. That sounds trivial — but it validates the method. The data surface a real, geographic relationship nobody hypothesized. The same process that finds "mountains" also finds "sick of" correlating with neuroticism and physical activity language correlating with emotional stability. The face-valid findings and the novel ones are products of the same process, which is the point. The paper is honest about limits. This is Facebook data, which carries sampling biases, demographic skews, and social desirability effects. People perform versions of themselves on social media, and what they perform is not identical to who they are in private. Some findings may not transfer to other text domains. Still, at the scale of seventy-five thousand participants and seven hundred million linguistic instances, with stringent multiple-testing correction applied throughout, the patterns that survive are real. What this study makes possible is not just description but discovery. A data-driven method at this scale can generate hypotheses — about personality, about life stages, about the relationship between activity and wellbeing — that structured theory would never produce. The practical reach is real: from text people already write, researchers can characterize psychological states with accuracy that surpasses what carefully designed instruments manage. That is the promise. The discomfort follows equally plainly. The same fingerprint that reveals neuroticism to a psychologist is present in the text itself, available to anyone with access to it. What Schwartz and colleagues have demonstrated is that our words are not merely communication. They are inadvertent autobiography — and it turns out they are surprisingly easy to read. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Seven hundred million words are sitting in a database. This is not a century of literature or a curated archive — just Facebook status updates from seventy-five thousand people, written over a few years without any thought of being studied. Schwartz and colleagues at the University of Pennsylvania let that data speak first. They imposed no categories, asked no leading questions, and what came back was a detailed portrait of who we are — our age, our gender, our deepest personality traits — written invisibly into the words we choose without ever knowing we were being read. To understand why this matters, you have to know what researchers had been doing before. The standard approach to language and personality is what Schwartz and colleagues call closed vocabulary. You start by asking humans to decide which words matter, sorting thousands of terms into fixed categories, and then counting how often those categories appear in someone's text. The best-known tool for this is Linguistic Inquiry and Word Count, or LIWC, a manually constructed lexicon with sixty-four categories covering everything from pronoun use to emotion words to family references. It's a serious, carefully built instrument. But it has a fundamental limitation: if you only look for the concepts you already put in the dictionary, you miss everything else.

The team cites a concrete example. The word "hug" correlates with agreeableness — but LIWC has no physical-affection category, so it simply vanishes from the analysis. Because the categories were built by human judges working from intuition, emergent patterns the data might reveal, such as a cluster of words tied to partying that distinguishes extraverts, won’t appear if nobody thought to include them. The numbers make this concrete. LIWC-based models predicted gender at seventy-eight point four percent accuracy and correlated with age at a correlation coefficient of zero point sixty-five. For personality, LIWC correlations clustered around zero point twenty-one to zero point twenty-nine across the Big Five traits. Those are not bad numbers. But the open-vocabulary approach beats all of them, sometimes substantially. The method Schwartz and colleagues call differential language analysis, or DLA, works in three steps. First, extract every word, phrase, and automatically generated topic from the text directly — no human curation. From fifteen point four million Facebook messages, that produced twenty-four thousand five hundred thirty word and phrase features, filtered to keep only those used by at least one percent of subjects.

On top of that, they ran Latent Dirichlet Allocation, or LDA, a topic-modeling technique that groups words into themes that the data itself reveals — not themes a researcher decided to look for. Second, test each feature for association with the target outcome using regression, controlling for covariates like age and gender so each result reflects a clean relationship. Because tens of thousands of features are being tested simultaneously, they apply strict Bonferroni correction, setting the significance threshold at a p-value below zero point zero zero one. Third, visualize the results with differential word clouds, where the size of each word reflects the strength of its correlation with the trait — not just how frequently it appears overall. The results for gender and age are the clearest demonstration of what the method can do. The sample had forty-six thousand four hundred twelve females and twenty-eight thousand two hundred forty-seven males. Using language features alone, the best predictive model reached ninety-one point nine percent out-of-sample accuracy. The previous best for language-only gender prediction was eighty-eight percent on Twitter. That jump comes from specific, readable signals. Females used more emotion words, more first-person singular pronouns, and more social expressions like "love you" and heart emoticons.

Males used more swear words and object references — "xbox" shows up as a significant predictor. One particularly striking micro-pattern: males were far more likely to precede a partner's title with the possessive "my." The phrases "my wife" and "my girlfriend" are strong male predictors. Females used "husband" and "boyfriend" without the possessive, or preceded them with other words — "amazing," "her" — so "my husband" simply didn't emerge as a parallel female signal. Age tells a story about life stages, not just vocabulary. The paper divided users into four groups — thirteen to eighteen, nineteen to twenty-two, twenty-three to twenty-nine, and thirty and older — and the linguistic shift across those groups tracks a recognizable arc: school to college to work to family. The youngest users wrote in emoticons and internet shorthand; the college cohort produced language around "semester," "drunk," and "hangover"; the twenty-something working group wrote about "at work" and "new job"; older users wrote about family. Across the full age range, the word "we" increases approximately linearly after age twenty-two while "I" monotonically decreases — a language-indexed shift toward social integration across the lifespan.

Personality is where the findings get genuinely revealing. The Big Five — extraversion, agreeableness, conscientiousness, neuroticism, and openness — are the standard framework in personality psychology, each representing a dimension along which people reliably differ. For each trait, DLA found specific, interpretable linguistic fingerprints. Neuroticism is the starkest. People who score high on neuroticism — emotional instability, anxiety, tendency toward negative affect — disproportionately use the phrase "sick of" and the word "depressed." Their language skews toward negative emotion and first-person singular pronouns, a pattern consistent with the self-focused, ruminative thinking that characterizes the trait. Openness to experience shows up in intellectual vocabulary, references to arts and culture, and a curiosity-inflected range of words that reflects engagement with ideas. Extraversion maps onto social language, positive emotions, and, notably, exclamation marks. Introverts, by contrast, show interest in topics like anime, Pokémon, and manga, along with Japanese emoticons. That kind of specificity is exactly what no closed lexicon would have predicted.

Conscientiousness brings words tied to schedules, achievement, and self-discipline — language that reflects organized, goal-directed behavior. Agreeableness maps onto warm, relational language, positivity directed outward rather than inward. One finding the team specifically flags as hypothesis-generating: words tied to physical activity and social participation — "basketball," "snowboarding," "church," "meetings" — correlate with emotional stability, the low end of neuroticism. The implication, carefully framed as a hypothesis rather than a conclusion, is that an active, socially embedded life may be linguistically and psychologically associated with stability. That is the kind of connection a theory-first approach wouldn't have surfaced because nobody thought to code a "physically active lifestyle" category into LIWC. The open-vocabulary model's predictive accuracy for personality consistently outperformed LIWC. LIWC correlations across the Big Five ranged from about zero point twenty-one to zero point twenty-nine; the best open-vocabulary model reached zero point forty-two for extraversion. For gender, the jump from seventy-eight point four percent to ninety-one point nine percent is the clearest summary of what the method adds. Some of the most interesting findings are the ones nobody would have thought to look for at all. People living in high elevations mention mountains. That sounds trivial — but it validates the method.

The data surface a real, geographic relationship nobody hypothesized. The same process that finds "mountains" also finds "sick of" correlating with neuroticism and physical activity language correlating with emotional stability. The face-valid findings and the novel ones are products of the same process, which is the point. The paper is honest about limits. This is Facebook data, which carries sampling biases, demographic skews, and social desirability effects. People perform versions of themselves on social media, and what they perform is not identical to who they are in private. Some findings may not transfer to other text domains. Still, at the scale of seventy-five thousand participants and seven hundred million linguistic instances, with stringent multiple-testing correction applied throughout, the patterns that survive are real. What this study makes possible is not just description but discovery. A data-driven method at this scale can generate hypotheses — about personality, about life stages, about the relationship between activity and wellbeing — that structured theory would never produce. The practical reach is real: from text people already write, researchers can characterize psychological states with accuracy that surpasses what carefully designed instruments manage. That is the promise. The discomfort follows equally plainly. The same fingerprint that reveals neuroticism to a psychologist is present in the text itself, available to anyone with access to it.

What Schwartz and colleagues have demonstrated is that our words are not merely communication. They are inadvertent autobiography — and it turns out they are surprisingly easy to read. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Computer Science