The Effects of Twitter Sentiment on Stock Price Returns

Gabriele Ranco, Darko Aleksovski, Guido Caldarelli, Miha Grćar, Igor MozetičView original
OverviewBalanceddamiaan voice
If you had to predict whether a stock was about to move up or down, and all you had was Twitter, would that be enough? Here's the honest answer from researchers who actually ran the numbers: mostly no. But at exactly the right moment — when the crowd's attention spikes — yes, and the signal lasts for days. The question has been circling financial research for over a decade. When social media exploded into daily life, it looked like a goldmine of behavioral data. Studies of search-engine queries linked daily volume to trading behavior. Research on news coverage tied media pessimism and earnings stories to price movements. Twitter specifically attracted a wave of attention: Bollen and colleagues found links between Twitter mood and aggregate market indices; other researchers showed that Twitter sentiment for a handful of retail companies had statistically significant relations with returns and volatility. A direct comparison study found that, at hourly resolution, sentiment carries more lead-time information than raw tweet volume alone. The field was building momentum. But most of this work examined aggregate indices or tiny company samples. Ranco, Aleksovski, Caldarelli, Grčar, and Mozetič wanted to do it rigorously, company by company, at scale. Their study covers the 30 companies in the Dow Jones Industrial Average — a natural, legible sample of household names — over 15 months from June 2013 to September 2014. They collected roughly 1.5 million tweets and built daily sentiment time series for each company. The classifier they trained is a supervised Support Vector Machine, a standard workhorse of text classification, but the annotation work behind it is substantial: over 100,000 manually labeled tweets, with more than 6,000 double-annotated so the team could measure how often two human raters agreed. That human agreement rate sets a ceiling — the best any automated system could reasonably hope for. The classifier used a two-step wrapper design to handle the three-way classification of positive, neutral, and negative sentiment as an ordinal problem, preserving the natural ordering rather than treating the three classes as arbitrary categories. Then they ran the obvious first tests. Over the full 15-month window, the Pearson correlation between daily Twitter sentiment polarity and daily stock returns is low. Individual company correlations range from about 0.12 for Travelers Companies to 0.36 for Caterpillar, with an aggregate summary correlation of 0.27. Those numbers sound moderate in isolation, but the signal here is weak, diluted by the many days when tweet volume is thin and sentiment is noisy. The Granger causality results are even more sobering. Granger causality asks whether knowing yesterday's Twitter sentiment actually improves your forecast of today's returns, beyond what the return series itself already tells you. It isn't about physical cause and effect — it's a predictive test. If past values of one series consistently add forecasting power for another series, beyond what that second series' own past provides, then the first is said to Granger-cause the second. Applied here: does yesterday's Twitter sentiment polarity help predict today's stock return? For sentiment polarity, the answer is barely. Only three of the 30 companies pass that test at the five percent significance level. Tweet volume fares better — it Granger-causes the absolute daily return, meaning volatility rather than direction, for roughly a third of companies, about ten out of thirty. So the crowd's attention predicts that something is about to happen. It just doesn't tell you, on average, which way. That flat, noisy baseline is exactly what motivates the paper's central move. Rather than treat every day as equivalent, the team asked: what happens if you focus only on the moments when Twitter is loudest? This is where they adapt the classic finance tool called an event study. In standard finance, an event study takes a known date — say, an earnings announcement — and measures whether returns around that date deviate from what a baseline model would predict. Those deviations are called abnormal returns, and their sum over a window of days is the cumulative abnormal return, or CAR. The market model that generates the baseline is simple to describe verbally: you regress each stock's return against the Dow Jones Industrial Average's return using a 120-day estimation window before the event, and whatever the model predicts as normal gets subtracted from the actual return. What's left is the abnormal part. Ranco and colleagues' innovation is to let Twitter itself define the events, automatically. For each company, they slide a window through time and compute the typical tweet volume as the median over that window. A day is flagged as a Twitter peak when the volume on that day is more than twice the baseline — a threshold they call phi-t equals two. Peaks within 21 days of each other get filtered out so events stay isolated. Over the 15-month sample, the algorithm found 260 peaks in total. Of those, 151 were earnings announcements, and the detector recovered 118 of them — a recall of 78 percent. After removing earnings announcement windows, 182 non-earnings events remained. Each detected peak gets assigned a prevailing sentiment score ranging from negative one to positive one. The team divides this into three classes: scores below 0.15 are negative, scores between 0.15 and 0.70 are neutral, and scores above 0.70 are positive. The finding is consistent and clean. The sign of aggregate cumulative abnormal returns after Twitter peaks tracks the prevailing sentiment on the peak day. Positive-sentiment peaks are followed by rising cumulative abnormal returns; negative-sentiment peaks are followed by falling ones. The magnitudes are modest — cumulative abnormal returns of about one to two percent on average. But modest doesn't mean unreal. For all detected events including earnings announcements, positive-sentiment peaks produce cumulative abnormal returns that are significant at the one percent level for ten days after the event. Negative-sentiment peaks produce losses roughly twice as large in absolute terms. Strip out the earnings announcements and look only at the organic, less predictable Twitter spikes: positive peaks still show cumulative abnormal returns significant at the one percent level for four days after the event, and negative peaks for eight days. Neutral sentiment peaks show little or only marginal significance. The direction of the crowd's mood, when the crowd is paying maximum attention, predicts the direction of price movement — and the effect persists across a trading week. This is the result that rescues the paper from the flat headline. The whole-period correlation analysis says: Twitter sentiment, averaged across all days, barely moves the needle. The event study says: that's because most days are noise. When you isolate the signal — the moments of peak collective attention — the sentiment carries genuine directional information, and the market follows. What the paper is careful not to claim is equally important. The one to two percent cumulative abnormal return figure is small. After trading costs, any strategy built on this signal would face serious headwinds. The authors flag that testing whether this finding survives as an actual trading strategy requires further work. They also note that the current approach treats all Twitter peaks as equivalent, ignoring what the tweets are actually about. Proposed next steps include automatic topic detection — methods like Latent Dirichlet Allocation, which identifies clusters of co-occurring words — to distinguish peaks driven by product launches, executive controversies, supply chain news, or just viral moments. The honest takeaway is not that social media moves markets in any simple, exploitable way. What Ranco and colleagues demonstrate is more specific and more interesting: Twitter sentiment, measured continuously across the full trading calendar, is a weak and largely unreliable signal. But crowd attention is not randomly distributed. When a stock captures a spike of public focus, the emotional tone of that conversation — whether predominantly positive or negative — carries information about the direction the stock will move over the following days. The market doesn't ignore what people are saying. It just takes a moment of unusual volume to make the signal audible above the noise. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

If you had to predict whether a stock was about to move up or down, and all you had was Twitter, would that be enough? Here's the honest answer from researchers who actually ran the numbers: mostly no. But at exactly the right moment — when the crowd's attention spikes — yes, and the signal lasts for days. The question has been circling financial research for over a decade. When social media exploded into daily life, it looked like a goldmine of behavioral data. Studies of search-engine queries linked daily volume to trading behavior. Research on news coverage tied media pessimism and earnings stories to price movements. Twitter specifically attracted a wave of attention: Bollen and colleagues found links between Twitter mood and aggregate market indices; other researchers showed that Twitter sentiment for a handful of retail companies had statistically significant relations with returns and volatility. A direct comparison study found that, at hourly resolution, sentiment carries more lead-time information than raw tweet volume alone. The field was building momentum. But most of this work examined aggregate indices or tiny company samples. Ranco, Aleksovski, Caldarelli, Grčar, and Mozetič wanted to do it rigorously, company by company, at scale.

Their study covers the 30 companies in the Dow Jones Industrial Average — a natural, legible sample of household names — over 15 months from June 2013 to September 2014. They collected roughly 1.5 million tweets and built daily sentiment time series for each company. The classifier they trained is a supervised Support Vector Machine, a standard workhorse of text classification, but the annotation work behind it is substantial: over 100,000 manually labeled tweets, with more than 6,000 double-annotated so the team could measure how often two human raters agreed. That human agreement rate sets a ceiling — the best any automated system could reasonably hope for. The classifier used a two-step wrapper design to handle the three-way classification of positive, neutral, and negative sentiment as an ordinal problem, preserving the natural ordering rather than treating the three classes as arbitrary categories. Then they ran the obvious first tests. Over the full 15-month window, the Pearson correlation between daily Twitter sentiment polarity and daily stock returns is low. Individual company correlations range from about 0.12 for Travelers Companies to 0.36 for Caterpillar, with an aggregate summary correlation of 0.27. Those numbers sound moderate in isolation, but the signal here is weak, diluted by the many days when tweet volume is thin and sentiment is noisy.

The Granger causality results are even more sobering. Granger causality asks whether knowing yesterday's Twitter sentiment actually improves your forecast of today's returns, beyond what the return series itself already tells you. It isn't about physical cause and effect — it's a predictive test. If past values of one series consistently add forecasting power for another series, beyond what that second series' own past provides, then the first is said to Granger-cause the second. Applied here: does yesterday's Twitter sentiment polarity help predict today's stock return? For sentiment polarity, the answer is barely. Only three of the 30 companies pass that test at the five percent significance level. Tweet volume fares better — it Granger-causes the absolute daily return, meaning volatility rather than direction, for roughly a third of companies, about ten out of thirty. So the crowd's attention predicts that something is about to happen. It just doesn't tell you, on average, which way. That flat, noisy baseline is exactly what motivates the paper's central move. Rather than treat every day as equivalent, the team asked: what happens if you focus only on the moments when Twitter is loudest?

This is where they adapt the classic finance tool called an event study. In standard finance, an event study takes a known date — say, an earnings announcement — and measures whether returns around that date deviate from what a baseline model would predict. Those deviations are called abnormal returns, and their sum over a window of days is the cumulative abnormal return, or CAR. The market model that generates the baseline is simple to describe verbally: you regress each stock's return against the Dow Jones Industrial Average's return using a 120-day estimation window before the event, and whatever the model predicts as normal gets subtracted from the actual return. What's left is the abnormal part. Ranco and colleagues' innovation is to let Twitter itself define the events, automatically. For each company, they slide a window through time and compute the typical tweet volume as the median over that window. A day is flagged as a Twitter peak when the volume on that day is more than twice the baseline — a threshold they call phi-t equals two. Peaks within 21 days of each other get filtered out so events stay isolated. Over the 15-month sample, the algorithm found 260 peaks in total. Of those, 151 were earnings announcements, and the detector recovered 118 of them — a recall of 78 percent. After removing earnings announcement windows, 182 non-earnings events remained.

Each detected peak gets assigned a prevailing sentiment score ranging from negative one to positive one. The team divides this into three classes: scores below 0.15 are negative, scores between 0.15 and 0.70 are neutral, and scores above 0.70 are positive. The finding is consistent and clean. The sign of aggregate cumulative abnormal returns after Twitter peaks tracks the prevailing sentiment on the peak day. Positive-sentiment peaks are followed by rising cumulative abnormal returns; negative-sentiment peaks are followed by falling ones. The magnitudes are modest — cumulative abnormal returns of about one to two percent on average. But modest doesn't mean unreal. For all detected events including earnings announcements, positive-sentiment peaks produce cumulative abnormal returns that are significant at the one percent level for ten days after the event. Negative-sentiment peaks produce losses roughly twice as large in absolute terms. Strip out the earnings announcements and look only at the organic, less predictable Twitter spikes: positive peaks still show cumulative abnormal returns significant at the one percent level for four days after the event, and negative peaks for eight days. Neutral sentiment peaks show little or only marginal significance. The direction of the crowd's mood, when the crowd is paying maximum attention, predicts the direction of price movement — and the effect persists across a trading week.

This is the result that rescues the paper from the flat headline. The whole-period correlation analysis says: Twitter sentiment, averaged across all days, barely moves the needle. The event study says: that's because most days are noise. When you isolate the signal — the moments of peak collective attention — the sentiment carries genuine directional information, and the market follows. What the paper is careful not to claim is equally important. The one to two percent cumulative abnormal return figure is small. After trading costs, any strategy built on this signal would face serious headwinds. The authors flag that testing whether this finding survives as an actual trading strategy requires further work. They also note that the current approach treats all Twitter peaks as equivalent, ignoring what the tweets are actually about. Proposed next steps include automatic topic detection — methods like Latent Dirichlet Allocation, which identifies clusters of co-occurring words — to distinguish peaks driven by product launches, executive controversies, supply chain news, or just viral moments. The honest takeaway is not that social media moves markets in any simple, exploitable way. What Ranco and colleagues demonstrate is more specific and more interesting: Twitter sentiment, measured continuously across the full trading calendar, is a weak and largely unreliable signal. But crowd attention is not randomly distributed.

When a stock captures a spike of public focus, the emotional tone of that conversation — whether predominantly positive or negative — carries information about the direction the stock will move over the following days. The market doesn't ignore what people are saying. It just takes a moment of unusual volume to make the signal audible above the noise. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Economics, Econometrics and Finance