Analysing How People Orient to and Spread Rumours in Social Media by Looking at Conversational Threads

Arkaitz Zubiaga, Maria Liakata, Rob Procter, Geraldine Wong Sak Hoi, Peter TolmieView original
OverviewBalancedalloy voice
When something breaks — a blast, a storm, or an unfolding attack — people instinctively turn to social media. That instinct arrives at exactly the wrong moment. Open networks let journalists and ordinary citizens report from the scene, sometimes ahead of mainstream outlets, but they also expose everyone to unverified claims that can travel far and fast before anyone knows whether they are true. Zubiaga and colleagues put this tension under a scientific microscope, and what they found should make anyone who reaches for Twitter during a crisis think twice. The word "rumour" in this research context is precise. The team defines a rumour as a circulating story of questionable veracity — apparently credible, hard to verify, and anxiety-producing enough that people want to know the actual truth. They anchor the stakes with two examples: a hacked Associated Press tweet in 2013 falsely reporting a White House attack that spooked financial markets, and the aftermath of Hurricane Sandy in 2012, when the Federal Emergency Management Agency had to build a dedicated rumour-control webpage for a city running on stressed mobile connections. These are not edge cases. They illustrate what the information environment looks like when breaking news and social media collide. The central question the paper poses is whether users behave differently around rumours that later prove true versus those that later prove false — and whether the difference is detectable in real time. To answer it, Zubiaga and colleagues built something genuinely careful. They used Twitter's streaming Application Programming Interface to collect tweets across nine newsworthy events: breaking stories like Ferguson, Charlie Hebdo, the Sydney siege, Ottawa, and Germanwings, plus four specific rumours tracked from the start, including the story of Putin missing and an Ebola claim about a soccer player. From those streams, highly retweeted candidate tweets were flagged, and here is what made the design distinctive: journalists embedded in the research team reviewed those candidates, labelled each source tweet as rumourous or not, grouped them into stories, and — critically — identified for each rumour the specific tweet that resolved it as true or false. Not a general sense that a rumour had been debunked, but a single, concrete resolving tweet within the timeline. The final annotated sample includes three hundred thirty rumour conversation threads containing four thousand eight hundred forty-two tweets. Within those, one hundred fifty-nine conversations were eventually confirmed true, sixty-eight false, and one hundred three remain unverified. For annotation, the team designed a three-dimensional scheme covering support versus denial, certainty, and evidentiality — whether a tweet attaches a URL, quotes a source, reports first-hand, or provides nothing. They crowdsourced the labelling through CrowdFlower, generating over sixty-eight thousand individual judgments from two hundred thirty-three annotators, with an overall agreement rate of sixty-two point three percent. Now to the findings. The most striking result about timing is this: true rumours are resolved far faster than false ones. The median true rumour is confirmed in about two hours from the source tweet. The median false rumour takes over fourteen hours to be debunked. The difference is highly significant — the paper reports a Wilcoxon signed-rank test with a p-value smaller than two point two times ten to the negative sixteenth power, which is as close to certainty as social science gets. Why does that gap matter? Because diffusion is front-loaded. Retweet activity in the dataset spikes hard in the first minutes after a tweet is posted, then fades within about twenty minutes. The tweets that generate the biggest early retweet bursts are pre-resolution tweets that support unverified rumours. A false rumour that takes fourteen hours to debunk spends the entirety of that high-velocity early window in an unresolved state — spreading before any corrective information can arrive. Per-event breakdowns confirm the scale: in the Ferguson data, twenty-six point seventy-eight percent of all retweets were of inaccurate content; in the Essien-Ebola story, that figure was sixty-nine point twenty-seven percent. The crowd's behaviour during that unresolved window is the second major finding, and it presents a portrait of systematic credulity. Zubiaga and colleagues measure what they call the support ratio: the number of supporting tweets divided by the combined total of supporting plus denying tweets. Before resolution, both true and false rumours have positive median support ratios. True rumours land at zero point sixteen, false ones at zero point zero seven — both on the supportive side, both reflecting a default "support first" posture while veracity is still unknown. The gap between them is statistically significant, but the key point is that neither type attracts more denial than support while unresolved. After a resolving tweet appears, the picture shifts. True rumours drop to a median support ratio of about zero point zero two. False rumours drop to minus zero point zero seven — the only condition in which denials reliably outnumber supports. In other words, correction happens, but only after the fact. The crowd is reactive, not predictive. And denying tweets, at any stage, receive the fewest retweets of any category. Two other annotated dimensions add texture. Certainty — how confident tweets sound — barely changes across the rumour lifecycle. For true rumours, the certainty median is zero point fifty-eight before resolution and zero point fifty-six after. The crowd sounds equally confident whether the claim is still up in the air or already confirmed. Evidentiality is more revealing: users attach more evidence while rumours are unverified, especially for false ones, with a median of zero point eighty-nine before debunking, dropping to zero point sixty-seven after. People are actively assembling and sharing evidence in support of claims that later turn out to be wrong. This connects directly to the paper's most uncomfortable finding, which involves the most trusted accounts. Zubiaga and colleagues operationalise user reputation using a follow ratio: the logarithm of followers divided by followees. Higher values mean a large audience relative to the accounts followed — a practical proxy for institutional reach. Of the forty-two users with a follow ratio of four or higher in the dataset, thirty-eight were news organisations. So the metric tracks real-world credibility well. The problem is that these highly reputable accounts are more likely to support rumours, to post with certainty, and to cite external evidence — regardless of whether the rumour they are supporting is actually true. Users with low follow ratios tend in the opposite direction: they show more denial, express more uncertainty, and provide less external citation. The signals people rely on to gauge trustworthiness — authoritativeness, confidence, and cited sources — are systematically uncorrelated with accuracy in this dataset. The paper links this to journalism's real-time pressures: professional norms push news organisations toward confident, evidence-accompanied posts, but the demand for speed competes with verification. The result is that an account's apparent credibility can amplify a false rumour more effectively than an anonymous tweet ever could. Taken together, the findings compose a coherent and troubling picture. False rumours take longer to resolve and therefore spend more time in the viral window. The crowd supports everything while it remains unverified, regardless of eventual truth value. Even the most reputable accounts fail to reliably distinguish true from false. Human collective judgment, at least in the form measurable here, does not function as a filter. That is what makes the paper's argument for automated assistance concrete rather than speculative. Because these behavioural patterns are consistent — support bias for unverified rumours, front-loaded retweet bursts, discrepant resolution times, and characteristic posting styles by user type — they are in principle learnable. Zubiaga and colleagues released their annotated dataset, with over sixty-eight thousand crowd judgments across four thousand eight hundred forty-two tweets, as a public resource. They argue that a useful automated system would need to classify support versus denial in real time, assess expressed certainty and evidentiality, and track conversational trajectories as they evolve toward or away from resolution. The opportunity is real because the patterns are real. The challenge is equally real: the signals are noisy, the time window is short, and the accounts that look most trustworthy are not the most accurate. What Zubiaga and colleagues have done is map the territory precisely enough that building something better is now a tractable problem rather than a hope. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

When something breaks — a blast, a storm, or an unfolding attack — people instinctively turn to social media. That instinct arrives at exactly the wrong moment. Open networks let journalists and ordinary citizens report from the scene, sometimes ahead of mainstream outlets, but they also expose everyone to unverified claims that can travel far and fast before anyone knows whether they are true.

Zubiaga and colleagues put this tension under a scientific microscope, and what they found should make anyone who reaches for Twitter during a crisis think twice.

The word "rumour" in this research context is precise. The team defines a rumour as a circulating story of questionable veracity — apparently credible, hard to verify, and anxiety-producing enough that people want to know the actual truth. They anchor the stakes with two examples: a hacked Associated Press tweet in 2013 falsely reporting a White House attack that spooked financial markets, and the aftermath of Hurricane Sandy in 2012, when the Federal Emergency Management Agency had to build a dedicated rumour-control webpage for a city running on stressed mobile connections.

These are not edge cases. They illustrate what the information environment looks like when breaking news and social media collide.

The central question the paper poses is whether users behave differently around rumours that later prove true versus those that later prove false — and whether the difference is detectable in real time.

To answer it, Zubiaga and colleagues built something genuinely careful. They used Twitter's streaming Application Programming Interface to collect tweets across nine newsworthy events: breaking stories like Ferguson, Charlie Hebdo, the Sydney siege, Ottawa, and Germanwings, plus four specific rumours tracked from the start, including the story of Putin missing and an Ebola claim about a soccer player. From those streams, highly retweeted candidate tweets were flagged, and here is what made the design distinctive: journalists embedded in the research team reviewed those candidates, labelled each source tweet as rumourous or not, grouped them into stories, and — critically — identified for each rumour the specific tweet that resolved it as true or false.

Not a general sense that a rumour had been debunked, but a single, concrete resolving tweet within the timeline.

The final annotated sample includes three hundred thirty rumour conversation threads containing four thousand eight hundred forty-two tweets. Within those, one hundred fifty-nine conversations were eventually confirmed true, sixty-eight false, and one hundred three remain unverified. For annotation, the team designed a three-dimensional scheme covering support versus denial, certainty, and evidentiality — whether a tweet attaches a URL, quotes a source, reports first-hand, or provides nothing.

They crowdsourced the labelling through CrowdFlower, generating over sixty-eight thousand individual judgments from two hundred thirty-three annotators, with an overall agreement rate of sixty-two point three percent.

Now to the findings. The most striking result about timing is this: true rumours are resolved far faster than false ones. The median true rumour is confirmed in about two hours from the source tweet.

The median false rumour takes over fourteen hours to be debunked. The difference is highly significant — the paper reports a Wilcoxon signed-rank test with a p-value smaller than two point two times ten to the negative sixteenth power, which is as close to certainty as social science gets.

Why does that gap matter? Because diffusion is front-loaded. Retweet activity in the dataset spikes hard in the first minutes after a tweet is posted, then fades within about twenty minutes.

The tweets that generate the biggest early retweet bursts are pre-resolution tweets that support unverified rumours. A false rumour that takes fourteen hours to debunk spends the entirety of that high-velocity early window in an unresolved state — spreading before any corrective information can arrive. Per-event breakdowns confirm the scale: in the Ferguson data, twenty-six point seventy-eight percent of all retweets were of inaccurate content; in the Essien-Ebola story, that figure was sixty-nine point twenty-seven percent.

The crowd's behaviour during that unresolved window is the second major finding, and it presents a portrait of systematic credulity. Zubiaga and colleagues measure what they call the support ratio: the number of supporting tweets divided by the combined total of supporting plus denying tweets. Before resolution, both true and false rumours have positive median support ratios.

True rumours land at zero point sixteen, false ones at zero point zero seven — both on the supportive side, both reflecting a default "support first" posture while veracity is still unknown. The gap between them is statistically significant, but the key point is that neither type attracts more denial than support while unresolved.

After a resolving tweet appears, the picture shifts. True rumours drop to a median support ratio of about zero point zero two. False rumours drop to minus zero point zero seven — the only condition in which denials reliably outnumber supports.

In other words, correction happens, but only after the fact. The crowd is reactive, not predictive. And denying tweets, at any stage, receive the fewest retweets of any category.

Two other annotated dimensions add texture. Certainty — how confident tweets sound — barely changes across the rumour lifecycle. For true rumours, the certainty median is zero point fifty-eight before resolution and zero point fifty-six after.

The crowd sounds equally confident whether the claim is still up in the air or already confirmed. Evidentiality is more revealing: users attach more evidence while rumours are unverified, especially for false ones, with a median of zero point eighty-nine before debunking, dropping to zero point sixty-seven after. People are actively assembling and sharing evidence in support of claims that later turn out to be wrong.

This connects directly to the paper's most uncomfortable finding, which involves the most trusted accounts. Zubiaga and colleagues operationalise user reputation using a follow ratio: the logarithm of followers divided by followees. Higher values mean a large audience relative to the accounts followed — a practical proxy for institutional reach.

Of the forty-two users with a follow ratio of four or higher in the dataset, thirty-eight were news organisations. So the metric tracks real-world credibility well.

The problem is that these highly reputable accounts are more likely to support rumours, to post with certainty, and to cite external evidence — regardless of whether the rumour they are supporting is actually true. Users with low follow ratios tend in the opposite direction: they show more denial, express more uncertainty, and provide less external citation. The signals people rely on to gauge trustworthiness — authoritativeness, confidence, and cited sources — are systematically uncorrelated with accuracy in this dataset.

The paper links this to journalism's real-time pressures: professional norms push news organisations toward confident, evidence-accompanied posts, but the demand for speed competes with verification. The result is that an account's apparent credibility can amplify a false rumour more effectively than an anonymous tweet ever could.

Taken together, the findings compose a coherent and troubling picture. False rumours take longer to resolve and therefore spend more time in the viral window. The crowd supports everything while it remains unverified, regardless of eventual truth value.

Even the most reputable accounts fail to reliably distinguish true from false. Human collective judgment, at least in the form measurable here, does not function as a filter.

That is what makes the paper's argument for automated assistance concrete rather than speculative. Because these behavioural patterns are consistent — support bias for unverified rumours, front-loaded retweet bursts, discrepant resolution times, and characteristic posting styles by user type — they are in principle learnable. Zubiaga and colleagues released their annotated dataset, with over sixty-eight thousand crowd judgments across four thousand eight hundred forty-two tweets, as a public resource.

They argue that a useful automated system would need to classify support versus denial in real time, assess expressed certainty and evidentiality, and track conversational trajectories as they evolve toward or away from resolution.

The opportunity is real because the patterns are real. The challenge is equally real: the signals are noisy, the time window is short, and the accounts that look most trustworthy are not the most accurate. What Zubiaga and colleagues have done is map the territory precisely enough that building something better is now a tractable problem rather than a hope.

This lecture was created by ennepō.

Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field.

Read when you can. Listen when you want to.

More in Social Sciences