Temporal Patterns of Happiness and Information in a Global Social NetworkHedonometrics and Twitter

Peter Sheridan Dodds, Kameron Decker Harris, Isabel M. Kloumann, Catherine A. Bliss, Christopher M. DanforthView original
OverviewBalanceddamiaan voice
Forty-six billion words are sitting in a database, typed by strangers on their phones. Somewhere in that pile is a signal — not about news, not about politics, but about how happy humanity was, hour by hour, for nearly three years. The question Dodds, Harris, Kloumann, Bliss, and Danforth set out to answer is whether you can actually measure collective human happiness from how people word a tweet. The problem they're addressing is older than Twitter. Happiness is a fundamental societal metric, but it's one that has always been awkward to capture. The standard approach is self-report: asking people how they’re doing and aggregating the answers into a survey index. The trouble is that surveys are slow, expensive, and retrospective. They tend to capture reflective well-being — how you evaluate your life when someone calls and asks — rather than experiential happiness, the feeling you're actually having right now. Meanwhile, economists have had gross domestic product on the table for decades. Gross domestic product is quantifiable, so it wins. The social sciences have known this is a problem. What they've lacked is a real-time instrument. Twitter, Dodds and colleagues argue, is different. It tends to generate in-the-moment expressions — not life evaluations but present-tense emotional output from millions of people simultaneously. If you can score the emotional content of that language, you have something close to a continuous, non-invasive, real-time sensor for collective mood. They call their instrument the hedonometer. Building it required solving a measurement problem first: which words carry happiness, and by how much? The team used Amazon's Mechanical Turk to survey human raters on a nine-point scale, from one for the saddest to nine for the happiest. They collected ratings for ten thousand two hundred twenty-two unique words — the labMT 1.0 word list — with fifty independent raters per word. Critically, the word list was not hand-selected for emotional content. The authors drew the top five thousand words from each of four large corpora — Twitter, Google Books, music lyrics, and the New York Times — and merged them. Frequency of use was the only criterion. That choice makes the metric defensible: it's not tuned to find emotion; it finds emotion because emotion is present in ordinary language. The scale anchors quickly once you see the examples. Laughter scores about eight point five. Food sits near seven point four four. Truck is nearly neutral at five point four eight. Hate falls to two point three four, funeral to two point one zero, and terrorist bottoms out at one point three zero. To compute the happiness score of a text, you take every scored word in that text, multiply each word's happiness score by how often it appears, sum those products, and divide by the total word count. You get a single number — a weighted average that reflects the relative prevalence of happier and sadder words. The instrument is also tunable. Words with average happiness scores close to the neutral midpoint of five carry little information about emotional direction, so the authors remove them. The width of that exclusion band is a parameter they call D-h. At a D-h of roughly one — excluding words scoring between four and six — three thousand six hundred eighty-six of the original ten thousand two hundred twenty-two words remain and together cover about twenty-two point seven percent of the Twitter corpus. For comparison, the older ANEW lexicon, the main prior tool for this kind of work, covers only about three point seven percent. That tenfold improvement in coverage is what makes the new instrument viable at scale. Now to the scale. The dataset runs from September two thousand eight through September two thousand eleven — thirty-three months, four point five eight six billion tweets, over forty-six billion words, from more than sixty-three million unique users. By August two thousand eleven, the collection rate had reached roughly twenty million tweets per day, about fourteen thousand per minute. The authors estimate they captured at least five percent of all tweets posted during the collection window. Their focused analyses used the period from May twenty-first, two thousand nine — when local timestamps became available — through December thirty-first, two thousand ten, giving them a clean window with time-of-day resolution. What does the hedonometer actually show? Three patterns dominate. First, a weekly cycle that is striking and stable. Saturdays score highest, averaging around six point zero six on the happiness scale, followed by Fridays and Sundays. Happiness then declines across the workweek to a nadir on Tuesday, near six point zero three. This ordering holds across four roughly equal time slices — it doesn’t belong to any single year or season. Second, a daily cycle. The happiest hour is five to six in the morning, with an average score near six point one two. Happiness falls steeply through the morning and afternoon and reaches its lowest point around ten to eleven at night, near six point zero two. That pattern is consistent, reproducible, and survives the removal of outlier dates with less than a tenth of a percent deviation. Third, and most dramatic, are outlier dates where events puncture the regular rhythm. Positive spikes cluster on holidays — Christmas, New Year's, Valentine's Day, Mother's Day, the Royal Wedding of April twenty-ninth, two thousand eleven. Negative crashes track collective trauma. The financial crisis bailout vote of September twenty-ninth, two thousand eight. Michael Jackson's death. The two thousand nine H1N1 pandemic onset. The March two thousand eleven Japan earthquake and tsunami. The killing of Osama Bin Laden on May second, two thousand eleven — which registers as the single lowest day across the entire three-year dataset. That last result deserves a pause. An event the American public broadly endorsed as justice produced the deepest collective verbal unhappiness in the dataset. What explains that? This is where Dodds and colleagues introduce word-shift graphs — a diagnostic tool that opens the hedonometer like a case file. A word-shift graph ranks every word by its percentage contribution to the change in average happiness between two texts. The top fifty words in a shift typically explain around sixty to seventy percent of the total change. On May second, two thousand eleven, the words surging in relative frequency included dead, death, killed, kill, and died. These are low-scoring words, and their sudden prevalence dragged the average down — regardless of the emotional valence people privately attached to the event. The hedonometer measures what words appear, not what people meant by them. Alongside happiness, the data carry a second signal: lexical richness. The authors quantify this with Simpson lexical size, which they define as the inverse of Simpson's concentration — the probability that two randomly chosen words from a text are identical. A higher Simpson lexical size means a more diverse, information-rich vocabulary. This measure also has a diurnal cycle. Lexical size peaks near five to six in the morning at around six hundred, then falls to a local minimum around nine to ten in the morning, rises modestly in early afternoon, and settles to a daily low near five hundred ten around ten to eleven at night. Early morning tweets are the richest in vocabulary. Late evening tweets are the most repetitive. The words driving that morning richness peak are illuminating. The words appearing less frequently at five to six in the morning — and therefore pulling lexical diversity down later in the day — include short function words and pronouns: I, a, the, de, me, que. The retweet marker RT is among the few words that increases later in the day, and its prevalence itself reduces diversity. The picture is of an early-morning Twitter that's compositionally richer, more original, and less templated. What matters most, though, is the relationship between the two signals: happiness and lexical richness are, in general, uncorrelated. The stream of tweets is not one trace but two independent rhythms layered on top of each other. The authors are direct about the instrument's limitations. The current hedonometer treats each word in isolation — it doesn't handle negation, so "not happy" counts as a sad-neutral pairing rather than an ironic positive. It doesn't score common phrases like "child abuse" or "sex scandal" as units. Language detection, Twitter sampling biases, and the risk of users gaming expressed sentiment are all acknowledged. These are genuine gaps. But the core result stands. Dodds and colleagues released labMT 1.0 as a public resource, along with the hedonometer methodology and the word-shift framework. Together they constitute a practical, transparent, frequency-based instrument for tracking experiential happiness at population scale in real time. Weekly cycles. Daily cycles. The fingerprints of collective trauma and collective joy, legible in the words people chose when they picked up their phones. Happiness has always been measurable, in principle. This paper built the meter. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Forty-six billion words are sitting in a database, typed by strangers on their phones. Somewhere in that pile is a signal — not about news, not about politics, but about how happy humanity was, hour by hour, for nearly three years. The question Dodds, Harris, Kloumann, Bliss, and Danforth set out to answer is whether you can actually measure collective human happiness from how people word a tweet. The problem they're addressing is older than Twitter. Happiness is a fundamental societal metric, but it's one that has always been awkward to capture. The standard approach is self-report: asking people how they’re doing and aggregating the answers into a survey index. The trouble is that surveys are slow, expensive, and retrospective. They tend to capture reflective well-being — how you evaluate your life when someone calls and asks — rather than experiential happiness, the feeling you're actually having right now. Meanwhile, economists have had gross domestic product on the table for decades. Gross domestic product is quantifiable, so it wins. The social sciences have known this is a problem. What they've lacked is a real-time instrument.

Twitter, Dodds and colleagues argue, is different. It tends to generate in-the-moment expressions — not life evaluations but present-tense emotional output from millions of people simultaneously. If you can score the emotional content of that language, you have something close to a continuous, non-invasive, real-time sensor for collective mood. They call their instrument the hedonometer. Building it required solving a measurement problem first: which words carry happiness, and by how much? The team used Amazon's Mechanical Turk to survey human raters on a nine-point scale, from one for the saddest to nine for the happiest. They collected ratings for ten thousand two hundred twenty-two unique words — the labMT 1.0 word list — with fifty independent raters per word. Critically, the word list was not hand-selected for emotional content. The authors drew the top five thousand words from each of four large corpora — Twitter, Google Books, music lyrics, and the New York Times — and merged them. Frequency of use was the only criterion. That choice makes the metric defensible: it's not tuned to find emotion; it finds emotion because emotion is present in ordinary language. The scale anchors quickly once you see the examples. Laughter scores about eight point five. Food sits near seven point four four.

Truck is nearly neutral at five point four eight. Hate falls to two point three four, funeral to two point one zero, and terrorist bottoms out at one point three zero. To compute the happiness score of a text, you take every scored word in that text, multiply each word's happiness score by how often it appears, sum those products, and divide by the total word count. You get a single number — a weighted average that reflects the relative prevalence of happier and sadder words. The instrument is also tunable. Words with average happiness scores close to the neutral midpoint of five carry little information about emotional direction, so the authors remove them. The width of that exclusion band is a parameter they call D-h. At a D-h of roughly one — excluding words scoring between four and six — three thousand six hundred eighty-six of the original ten thousand two hundred twenty-two words remain and together cover about twenty-two point seven percent of the Twitter corpus. For comparison, the older ANEW lexicon, the main prior tool for this kind of work, covers only about three point seven percent. That tenfold improvement in coverage is what makes the new instrument viable at scale.

Now to the scale. The dataset runs from September two thousand eight through September two thousand eleven — thirty-three months, four point five eight six billion tweets, over forty-six billion words, from more than sixty-three million unique users. By August two thousand eleven, the collection rate had reached roughly twenty million tweets per day, about fourteen thousand per minute. The authors estimate they captured at least five percent of all tweets posted during the collection window. Their focused analyses used the period from May twenty-first, two thousand nine — when local timestamps became available — through December thirty-first, two thousand ten, giving them a clean window with time-of-day resolution. What does the hedonometer actually show? Three patterns dominate. First, a weekly cycle that is striking and stable. Saturdays score highest, averaging around six point zero six on the happiness scale, followed by Fridays and Sundays. Happiness then declines across the workweek to a nadir on Tuesday, near six point zero three. This ordering holds across four roughly equal time slices — it doesn’t belong to any single year or season. Second, a daily cycle. The happiest hour is five to six in the morning, with an average score near six point one two. Happiness falls steeply through the morning and afternoon and reaches its lowest point around ten to eleven at night, near six point zero two.

That pattern is consistent, reproducible, and survives the removal of outlier dates with less than a tenth of a percent deviation. Third, and most dramatic, are outlier dates where events puncture the regular rhythm. Positive spikes cluster on holidays — Christmas, New Year's, Valentine's Day, Mother's Day, the Royal Wedding of April twenty-ninth, two thousand eleven. Negative crashes track collective trauma. The financial crisis bailout vote of September twenty-ninth, two thousand eight. Michael Jackson's death. The two thousand nine H1N1 pandemic onset. The March two thousand eleven Japan earthquake and tsunami. The killing of Osama Bin Laden on May second, two thousand eleven — which registers as the single lowest day across the entire three-year dataset. That last result deserves a pause. An event the American public broadly endorsed as justice produced the deepest collective verbal unhappiness in the dataset. What explains that? This is where Dodds and colleagues introduce word-shift graphs — a diagnostic tool that opens the hedonometer like a case file. A word-shift graph ranks every word by its percentage contribution to the change in average happiness between two texts. The top fifty words in a shift typically explain around sixty to seventy percent of the total change.

On May second, two thousand eleven, the words surging in relative frequency included dead, death, killed, kill, and died. These are low-scoring words, and their sudden prevalence dragged the average down — regardless of the emotional valence people privately attached to the event. The hedonometer measures what words appear, not what people meant by them. Alongside happiness, the data carry a second signal: lexical richness. The authors quantify this with Simpson lexical size, which they define as the inverse of Simpson's concentration — the probability that two randomly chosen words from a text are identical. A higher Simpson lexical size means a more diverse, information-rich vocabulary. This measure also has a diurnal cycle. Lexical size peaks near five to six in the morning at around six hundred, then falls to a local minimum around nine to ten in the morning, rises modestly in early afternoon, and settles to a daily low near five hundred ten around ten to eleven at night. Early morning tweets are the richest in vocabulary. Late evening tweets are the most repetitive. The words driving that morning richness peak are illuminating. The words appearing less frequently at five to six in the morning — and therefore pulling lexical diversity down later in the day — include short function words and pronouns: I, a, the, de, me, que. The retweet marker RT is among the few words that increases later in the day, and its prevalence itself reduces diversity.

The picture is of an early-morning Twitter that's compositionally richer, more original, and less templated. What matters most, though, is the relationship between the two signals: happiness and lexical richness are, in general, uncorrelated. The stream of tweets is not one trace but two independent rhythms layered on top of each other. The authors are direct about the instrument's limitations. The current hedonometer treats each word in isolation — it doesn't handle negation, so "not happy" counts as a sad-neutral pairing rather than an ironic positive. It doesn't score common phrases like "child abuse" or "sex scandal" as units. Language detection, Twitter sampling biases, and the risk of users gaming expressed sentiment are all acknowledged. These are genuine gaps. But the core result stands. Dodds and colleagues released labMT 1.0 as a public resource, along with the hedonometer methodology and the word-shift framework. Together they constitute a practical, transparent, frequency-based instrument for tracking experiential happiness at population scale in real time. Weekly cycles. Daily cycles. The fingerprints of collective trauma and collective joy, legible in the words people chose when they picked up their phones. Happiness has always been measurable, in principle. This paper built the meter. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Physics and Astronomy