Coronavirus Goes ViralQuantifying the COVID-19 Misinformation Epidemic on Twitter

Ramez Kouzy, Joseph Abi Jaoude, Afif Kraitem, Molly B. El Alam, Basil S. Karam, Elio Adib, Jabra Zarka, Cindy Traboulsi, Elie A. Akl, Khalil BaddourView original
OverviewBalancedalloy voice
If you think back to those first weeks of COVID, it wasn't just a virus moving fast. It was claims, screenshots, threads, and advice from your cousin's roommate's doctor. Kouzy and colleagues have a simple name for that surge: an infodemic—information washing over us faster than we could check it. Inside that flood, they focused on one slippery piece: medical misinformation, which they define as a factual claim that's currently false because there's no scientific support for it. No peer review, no professional verification, just a message that sounds confident enough to spread. In a pandemic, that's not a nuisance. It's a public health hazard. Their goal was practical: put numbers on what was flying around Twitter early in the outbreak, and see which kinds of posts were most likely to be wrong or uncheckable. They weren't trying to shame individual users, but to map the terrain so health agencies and platforms could move faster next time. That framing echoes what we saw during Ebola and Zika—social media can amplify fear and rumor when uncertainty is high. Here, they wanted to turn the hand-waving into a measurement. So let's talk about how they did it. On a single day—February twenty-seventh, two thousand twenty—they pulled tweets using the Twitter Archiver add-on. The search used fourteen trending terms identified by Symplur, a health-care analytics group: eleven hashtags and three keywords tied to COVID. They kept it to English and, crucially, only included tweets that had already drawn at least five retweets. The idea was to focus on content with some traction and sidestep the noise of a single unamplified post. From that stream, they analyzed six hundred seventy-three tweets. It's a snapshot, not a census, but it's early enough to capture the mix of fear, curiosity, and speculation that defined that week. What they did next matters as much as the sample. Every tweet came with a profile: what type of account posted it, whether Twitter had verified it as an authentic account of public interest, the tone the tweet used, and what domain it touched—medical and public health, finance, politics. Then they coded the content itself. Claims were cross-checked against sources like the World Health Organization, the Centers for Disease Control and Prevention, peer-reviewed journals, and major news outlets. If the claim could be clearly refuted by those references, it was misinformation. If the claim couldn't be proven true or false by those references—too vague, too early, or simply not addressed—it was labeled unverifiable. That coding choice is important. It separates "wrong" from "not yet knowable." Now to the shape of the conversation. The loudest voices in this sample weren't institutions. They were informal personal or group accounts, responsible for about two-thirds of the tweets. Only about a fifth of accounts were verified. The vast majority of tweets were serious in tone. This wasn't mostly jokes; it was people trying to share things they believed mattered. Most of the content focused on medical and public health topics. Inside that serious corpus, roughly a quarter of tweets contained misinformation and about a sixth were unverifiable. That's the headline. One in four was wrong; another chunk couldn't be checked by the best references available at the time. Where did the problems cluster? Start with account type. Informal accounts—the ones not tied to a news organization, a government, or a health agency—had the highest rate of misinformation. About a third of their tweets were coded as misinformation, compared with around fifteen percent for everyone else. That's a big gap, and it fits a broader pattern: institutional accounts, with editorial oversight or professional stakes, tended to be more accurate. Health care and public health accounts were standouts on the low end. Their misinformation rate was about twelve percent, far below the informal group. Business, nongovernmental, and government accounts also ran relatively low, and even news outlets and journalists, while not perfect, hovered closer to the bottom than the top. In other words, who you are, or who your account represents, signaled something real about the likelihood you were posting a false claim. Verification status told a similar story, and it was stark. Verified accounts—Twitter's label for authentic accounts of public interest—posted misinformation at a rate of about thirteen percent. Unverified accounts? Roughly thirty-one percent. That's not a tiny difference; that's a doubling. It held even when you looked at audience size. Accounts with more than about eleven thousand followers had a lower misinformation rate than smaller accounts—around twenty percent versus thirty-four percent. That threshold may seem oddly precise, but it came from the data; the median follower count split the sample into two roughly equal groups. The takeaway isn't that eleven thousand forty-five is a magical number. It's that bigger, more established audiences were less likely to push false claims in that moment. Here's the twist that matters for anyone chasing virality: likes and retweets did not separate truth from falsehood. The rate of misinformation did not change with the number of likes a tweet got. Same story for retweets. Those p-values were zero point ninety-eight for likes and zero point thirty-six for retweets. They were not close. In plain language, a bad claim was just as likely to get engagement as a good one in this sample. That's unsettling if you use engagement as a proxy for quality, which, let's be honest, is how most of us scroll. Unverifiable information followed its own logic. Verified accounts again did better, with an unverifiable rate under ten percent, while unverified accounts landed much higher. But follower count didn't cleanly predict unverifiable content the way it did for outright misinformation. That difference makes intuitive sense. "Not yet knowable" isn't about sloppiness or intent; it's about the state of knowledge at a point in time. A large audience might still ask questions or pass along anecdotes that can't be checked that week, even if they're more careful about claims that can be debunked. Even the words people chose in their tweets mattered. Some terms were magnets for trouble. The hashtag "#2019_nCov" was associated with the highest rate of misinformation in the set. Meanwhile, tweets that used "COVID-19" or "#nCov19" were among the least likely to contain misinformation. For unverifiable content, the catchall "Corona" popped up with the highest rate, whereas "COVID-19" and "#coronavirusoutbreak" sat on the other end, with the lowest. This isn't a pedantic point about nomenclature. Terms that aligned with formal naming conventions mapped, in this snapshot, to more reliable content. Vaguer or earlier labels pulled in more noise. Behind these findings is a pretty straightforward statistics toolkit. The team summarized who posted what and then used chi-square tests to see whether misinformation or unverifiable content was more common in one group than another, setting statistical significance at the standard two-sided zero point zero five. They did the work in a standard analysis package. Nothing fancy. The credibility came from the coding discipline—defining misinformation against external references—and from separating out claims that were simply not testable that week. Before we run too far with the numbers, it's worth staying honest about what this study could and couldn't say. It's English only. It's one day. It's posts that had already earned at least five retweets, which means it intentionally studied content with a hint of momentum and left the long tail alone. It relies on manual coding, even if the team tried to anchor every judgment to the World Health Organization, the Centers for Disease Control and Prevention, journals, and major outlets. And it's small—six hundred seventy-three tweets is a tractable sample, not the firehose. All of that narrows generalizability. The authors are clear about those boundaries. Still, some patterns feel durable. The absence of any link between engagement and accuracy is a big one. If likes and retweets don't filter out misinformation in a crisis, then our usual instinct to trust what seems popular can backfire. The gap between informal and institutional voices also tracks with what many of us sensed: expertise and accountability matter, even in the messiness of real-time science. Verification status wasn't just a badge—it correlated with fewer false and fewer uncheckable claims. You can probably hear the policy levers rattling. Kouzy and colleagues argue that counter-messaging has to start early and come from health care and public health accounts that audiences can find and trust. Platforms can help by amplifying credible sources and making verification cues more visible when uncertainty spikes. Even the plumbing of hashtags and keywords matters. Standardizing around accurate terms—think "COVID-19" rather than looser labels—may reduce the surface area for rumor to spread, because those streams, at least in this sample, carried cleaner information. They also call for collaboration. Physicians, medical associations, journals, public health agencies, and newsrooms can coordinate faster when the next infodemic begins, and replace bad claims with vetted answers rather than just calling them out. None of this requires draconian censorship to be effective. It's about making the trustworthy stuff easier to find, earlier, and across more of the information network that people actually use. Let's end where we began, with what these numbers mean when you're just scrolling on your phone. In this early COVID snapshot, about one in four tweets with some traction was wrong, and a sizable chunk couldn't be checked. The people most likely to be accurate were the ones with expertise, verification, or institutional roles. Popularity didn't tell you truth from falsehood. Words mattered. Speed cut both ways—it helped good information spread, and it gave bad claims a head start. That's not an indictment of the crowd. It's a map of the terrain you were walking through. So the next time the world tilts and your feed lights up, a few habits might help. Look for the names that carry responsibility. Notice the terms. Treat virality as a signal of reach, not of reliability. And remember that in a true emergency, the difference between "we don't know yet" and "this is false" is not academic. It's the space where public health lives. Kouzy's team handed us a ruler for that space. It's on all of us—platforms, pros, and the rest of us with thumbs—to use it.

If you think back to those first weeks of COVID, it wasn't just a virus moving fast. It was claims, screenshots, threads, and advice from your cousin's roommate's doctor. Kouzy and colleagues have a simple name for that surge: an infodemic—information washing over us faster than we could check it.

Inside that flood, they focused on one slippery piece: medical misinformation, which they define as a factual claim that's currently false because there's no scientific support for it. No peer review, no professional verification, just a message that sounds confident enough to spread. In a pandemic, that's not a nuisance. It's a public health hazard.

Their goal was practical: put numbers on what was flying around Twitter early in the outbreak, and see which kinds of posts were most likely to be wrong or uncheckable. They weren't trying to shame individual users, but to map the terrain so health agencies and platforms could move faster next time. That framing echoes what we saw during Ebola and Zika—social media can amplify fear and rumor when uncertainty is high. Here, they wanted to turn the hand-waving into a measurement.

So let's talk about how they did it. On a single day—February twenty-seventh, two thousand twenty—they pulled tweets using the Twitter Archiver add-on. The search used fourteen trending terms identified by Symplur, a health-care analytics group: eleven hashtags and three keywords tied to COVID.

They kept it to English and, crucially, only included tweets that had already drawn at least five retweets. The idea was to focus on content with some traction and sidestep the noise of a single unamplified post. From that stream, they analyzed six hundred seventy-three tweets.

It's a snapshot, not a census, but it's early enough to capture the mix of fear, curiosity, and speculation that defined that week.

What they did next matters as much as the sample. Every tweet came with a profile: what type of account posted it, whether Twitter had verified it as an authentic account of public interest, the tone the tweet used, and what domain it touched—medical and public health, finance, politics. Then they coded the content itself.

Claims were cross-checked against sources like the World Health Organization, the Centers for Disease Control and Prevention, peer-reviewed journals, and major news outlets. If the claim could be clearly refuted by those references, it was misinformation. If the claim couldn't be proven true or false by those references—too vague, too early, or simply not addressed—it was labeled unverifiable. That coding choice is important. It separates "wrong" from "not yet knowable."

Now to the shape of the conversation. The loudest voices in this sample weren't institutions. They were informal personal or group accounts, responsible for about two-thirds of the tweets.

Only about a fifth of accounts were verified. The vast majority of tweets were serious in tone. This wasn't mostly jokes; it was people trying to share things they believed mattered.

Most of the content focused on medical and public health topics. Inside that serious corpus, roughly a quarter of tweets contained misinformation and about a sixth were unverifiable. That's the headline.

One in four was wrong; another chunk couldn't be checked by the best references available at the time.

Where did the problems cluster? Start with account type. Informal accounts—the ones not tied to a news organization, a government, or a health agency—had the highest rate of misinformation.

About a third of their tweets were coded as misinformation, compared with around fifteen percent for everyone else. That's a big gap, and it fits a broader pattern: institutional accounts, with editorial oversight or professional stakes, tended to be more accurate. Health care and public health accounts were standouts on the low end.

Their misinformation rate was about twelve percent, far below the informal group. Business, nongovernmental, and government accounts also ran relatively low, and even news outlets and journalists, while not perfect, hovered closer to the bottom than the top. In other words, who you are, or who your account represents, signaled something real about the likelihood you were posting a false claim.

Verification status told a similar story, and it was stark. Verified accounts—Twitter's label for authentic accounts of public interest—posted misinformation at a rate of about thirteen percent. Unverified accounts?

Roughly thirty-one percent. That's not a tiny difference; that's a doubling. It held even when you looked at audience size.

Accounts with more than about eleven thousand followers had a lower misinformation rate than smaller accounts—around twenty percent versus thirty-four percent. That threshold may seem oddly precise, but it came from the data; the median follower count split the sample into two roughly equal groups. The takeaway isn't that eleven thousand forty-five is a magical number.

It's that bigger, more established audiences were less likely to push false claims in that moment.

Here's the twist that matters for anyone chasing virality: likes and retweets did not separate truth from falsehood. The rate of misinformation did not change with the number of likes a tweet got. Same story for retweets.

Those p-values were zero point ninety-eight for likes and zero point thirty-six for retweets. They were not close. In plain language, a bad claim was just as likely to get engagement as a good one in this sample.

That's unsettling if you use engagement as a proxy for quality, which, let's be honest, is how most of us scroll.

Unverifiable information followed its own logic. Verified accounts again did better, with an unverifiable rate under ten percent, while unverified accounts landed much higher. But follower count didn't cleanly predict unverifiable content the way it did for outright misinformation.

That difference makes intuitive sense. "Not yet knowable" isn't about sloppiness or intent; it's about the state of knowledge at a point in time. A large audience might still ask questions or pass along anecdotes that can't be checked that week, even if they're more careful about claims that can be debunked.

Even the words people chose in their tweets mattered. Some terms were magnets for trouble. The hashtag "#2019_nCov" was associated with the highest rate of misinformation in the set.

Meanwhile, tweets that used "COVID-19" or "#nCov19" were among the least likely to contain misinformation. For unverifiable content, the catchall "Corona" popped up with the highest rate, whereas "COVID-19" and "#coronavirusoutbreak" sat on the other end, with the lowest. This isn't a pedantic point about nomenclature.

Terms that aligned with formal naming conventions mapped, in this snapshot, to more reliable content. Vaguer or earlier labels pulled in more noise.

Behind these findings is a pretty straightforward statistics toolkit. The team summarized who posted what and then used chi-square tests to see whether misinformation or unverifiable content was more common in one group than another, setting statistical significance at the standard two-sided zero point zero five. They did the work in a standard analysis package.

Nothing fancy. The credibility came from the coding discipline—defining misinformation against external references—and from separating out claims that were simply not testable that week.

Before we run too far with the numbers, it's worth staying honest about what this study could and couldn't say. It's English only. It's one day.

It's posts that had already earned at least five retweets, which means it intentionally studied content with a hint of momentum and left the long tail alone. It relies on manual coding, even if the team tried to anchor every judgment to the World Health Organization, the Centers for Disease Control and Prevention, journals, and major outlets. And it's small—six hundred seventy-three tweets is a tractable sample, not the firehose.

All of that narrows generalizability. The authors are clear about those boundaries.

Still, some patterns feel durable. The absence of any link between engagement and accuracy is a big one. If likes and retweets don't filter out misinformation in a crisis, then our usual instinct to trust what seems popular can backfire.

The gap between informal and institutional voices also tracks with what many of us sensed: expertise and accountability matter, even in the messiness of real-time science. Verification status wasn't just a badge—it correlated with fewer false and fewer uncheckable claims.

You can probably hear the policy levers rattling. Kouzy and colleagues argue that counter-messaging has to start early and come from health care and public health accounts that audiences can find and trust. Platforms can help by amplifying credible sources and making verification cues more visible when uncertainty spikes.

Even the plumbing of hashtags and keywords matters. Standardizing around accurate terms—think "COVID-19" rather than looser labels—may reduce the surface area for rumor to spread, because those streams, at least in this sample, carried cleaner information.

They also call for collaboration. Physicians, medical associations, journals, public health agencies, and newsrooms can coordinate faster when the next infodemic begins, and replace bad claims with vetted answers rather than just calling them out. None of this requires draconian censorship to be effective.

It's about making the trustworthy stuff easier to find, earlier, and across more of the information network that people actually use.

Let's end where we began, with what these numbers mean when you're just scrolling on your phone. In this early COVID snapshot, about one in four tweets with some traction was wrong, and a sizable chunk couldn't be checked. The people most likely to be accurate were the ones with expertise, verification, or institutional roles.

Popularity didn't tell you truth from falsehood. Words mattered. Speed cut both ways—it helped good information spread, and it gave bad claims a head start.

That's not an indictment of the crowd. It's a map of the terrain you were walking through.

So the next time the world tilts and your feed lights up, a few habits might help. Look for the names that carry responsibility. Notice the terms.

Treat virality as a signal of reach, not of reliability. And remember that in a true emergency, the difference between "we don't know yet" and "this is false" is not academic. It's the space where public health lives.

Kouzy's team handed us a ruler for that space. It's on all of us—platforms, pros, and the rest of us with thumbs—to use it.

More in Social Sciences