Do Altmetrics Work? Twitter and Ten Other Social Web Services
A tweet arrives. Someone has shared a link to a scientific paper. Somewhere, a number ticks up. That number — that single social web mention — is now being used by universities and funding bodies as evidence that the research matters. So here is the question this episode answers flat: does it actually mean anything? Citations have long been the gold standard of scholarly impact because they explicitly record when one piece of research is used by another. But citations take years to accumulate. That delay creates pressure to find faster alternatives. Enter altmetrics: counts of mentions and interactions across the social web — tweets, Facebook wall posts, blog coverage, mainstream media mentions, forum posts, and more. Publishers and tracking services like altmetric.com already use these signals for rapid promotion. Tweets about a paper typically peak on its publication day; blog coverage usually appears within a month. That speed is the whole appeal. But Thelwall, Haustein, Larivière, and Sugimoto asked the harder question — not whether altmetrics are fast, but whether they're valid.
Their study tested eleven altmetric sources against Web of Science citation counts for PubMed-indexed articles published between 2010 and 2012. The eleven platforms were Twitter, Facebook wall posts, research highlights from Nature Publishing Group, blogs drawn from a curated list of around two thousand two hundred science blogs, mainstream media from a curated list of roughly sixty outlets, Google+, LinkedIn, Pinterest, Reddit, forum posts from two scraped forums, and question-and-answer sites including Stack Exchange. Depending on the platform, sample sizes ranged from one hundred eighty-two articles up to two hundred eight thousand seven hundred thirty-nine, and the study covered as many as one thousand eight hundred ninety-one journals per metric. That's a genuinely large-scale comparison. The methodological challenge was real. Raw correlations between altmetric scores and citation counts are badly distorted by time. Social media attention spikes fast; citations accumulate slowly. So a newer article might have lots of tweets but few citations simply because it hasn't had time to accumulate them — not because it's weaker science. To get around this, Thelwall and colleagues introduced a sign test. For each journal and each metric, they ordered articles by their PubMed identifier and compared each article only to the two published immediately before and after it in the same journal.
The question becomes simple: if this article has a higher altmetric score than the average of its two neighbors, does it also tend to have more citations? Outcomes were classified as successes or failures. This design neutralizes age differences and field effects in one move, and any consistent excess of successes over failures becomes direct evidence of a real association. The results were mixed, but in an instructive way. Six platforms showed statistically significant associations between higher altmetric scores and higher citations: Twitter, Facebook wall posts, research highlights, blogs, mainstream media, and forums. The numbers from the sign test make the pattern tangible. For Twitter, there were twenty-four thousand three hundred fifteen successes versus eighteen thousand five hundred seventy-six failures — that's fifty-seven percent versus forty-three percent, across more than two hundred thousand tests. For blogs, it was sixty percent successes to forty percent failures. Forum posts showed the most striking split: eighty-six percent successes to fourteen percent failures, though the raw numbers there were small — just nineteen successes against three failures.
Google+ was the notable exception. Despite having articles tagged, the results were not statistically significant: four hundred twenty-six successes versus three hundred seventy-eight failures, barely above chance. The authors flag this as possibly a statistical anomaly rather than a genuine difference in how Google+ behaves. For LinkedIn, Pinterest, Reddit, and question-and-answer sites, the evidence was simply insufficient. These platforms had too few articles with non-zero scores to support any conclusion. The data were sparse enough that drawing inferences would be unreliable. It's worth being precise about what "association" means here. Thelwall and colleagues explicitly did not test predictive power — they did not ask whether high altmetric scores today predict high citations tomorrow. The finding is correlational: articles that get more social attention on the six validated platforms also tend to be cited more. That's not nothing. But it's not causation, and it's not a forecasting tool. Now here's the part that tends to get buried. The relationship between altmetrics and citations can disappear entirely — or reverse — depending on when you compare articles. This is the time problem.
A raw cross-sectional correlation for Twitter across the whole dataset actually came out negative. When the team restricted the analysis to two thousand ten articles only, they got a small but significantly negative correlation of negative zero point two four. When they removed the influence of time using a partial correlation controlling for PubMed identifier, that number moved to essentially zero — zero point zero zero nine. The same data, the same metric, three different answers depending on how you handle time. That's not a technicality. That's a warning. If you compare a paper published in December against one published in January using raw citation counts, you're not comparing scientific quality — you're comparing age. The December paper will look worse simply because it's newer. Thelwall and colleagues are direct about this: time from publication must be considered when using altmetrics to rank or evaluate articles. Publishers and scientometricians who rank articles by unadjusted altmetric-citation comparisons will systematically disadvantage recent work. The second major limitation is coverage. Even for Twitter — the best-performing metric by far — only a fraction of published articles receive any mention at all. For every other platform, coverage is below twenty percent of articles, and in many cases substantially below that.
The research highlights metric covered only twenty percent as many articles as Twitter. LinkedIn-flagged articles represented just zero point zero four percent as many as tweeted articles. Most papers, on most platforms, have a score of zero. This creates a specific problem for interpretation. The study's sample came from altmetric.com and included only articles with at least one non-zero altmetric mention. The team explicitly excluded zero-score articles from their conclusions because a missing record and a genuine zero look identical in the data. So all of these findings — the six validated associations, the sign-test results — apply only to articles that attracted some social attention in the first place. The majority of published research is invisible on most of these platforms, and this study says nothing about that majority. The practical implication follows directly. Altmetrics can flag the occasional exceptional or above-average article. They cannot serve as a general-purpose substitute for citation counts, because they're simply not present for most of the literature. A metric that's silent on eighty percent or more of published work can highlight outliers, but it can't discriminate among typical articles.
So where does this leave us? Thelwall and colleagues establish a careful, empirically grounded baseline. Six platforms — Twitter, Facebook wall posts, research highlights, blogs, mainstream media, and forums — show validated associations with citation counts in medical and biological sciences, for articles that have at least one altmetric mention. That's a meaningful finding. It means these numbers aren't random noise. For the papers that get noticed on these platforms, the attention does tend to track with how much other researchers eventually cite the work. But the study also draws clear boundaries around that finding. The magnitude of any correlation was not estimated — the sign test tells you direction, not strength. Coverage is low everywhere except possibly Twitter. And publication timing can flip the apparent relationship entirely, which means any ranking or evaluation that uses altmetrics without controlling for article age is working with a distorted picture. If someone is going to use a tweet count to evaluate a scientist's work — and increasingly, people are — they should at minimum know which platforms have been empirically validated, which haven't, and why comparing a paper from January to one from December without an age adjustment can reverse the conclusion entirely. That's what this study gives us: not a verdict on altmetrics, but the minimum standard for using them responsibly. This lecture was created by ennepō.
Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- ‘Predatory’ open access: a longitudinal study of article volumes and market characteristics
- Do Pressures to Publish Increase Scientists' Bias? An Empirical Support from US States Data
- A Dual-Self Model of Impulse Control
- Self-Selected or Mandated, Open Access Increases Citation Impact for Higher Quality Research
- Citation Advantage of Open Access Articles
- The heterogeneity statistic I2 can be biased in small meta-analyses