Citation Advantage of Open Access Articles

Günther EysenbachView original
OverviewBalancedharper voice
Open access articles get cited more. Not just a little more — in one carefully controlled study, they were nearly three times more likely to be cited by the end of their first year. Now hold that thought for a moment, because the real question isn't the number. It's whether that advantage is real or whether it's just a shadow cast by something else entirely. That something else has a name in the literature: quality bias. The concern is simple and hard to dismiss. Perhaps scientists make their best work freely available because a high-profile result attracts reprint requests, or because a professor puts a landmark paper on a course website, or because it only takes one enthusiastic co-author to upload a preprint. If that's the case, then the association between open access and citations tells you nothing about access at all. It tells you that good papers get read. The question Günther Eysenbach set out to answer is whether open access does anything on its own, stripped of that confounding factor. Every earlier study on this question was cross-sectional — a snapshot comparing citation counts between papers that happened to be freely available and those that weren't, with no ability to control for what made those papers different in the first place. You couldn't tell whether access caused citations or citations caused access. Eysenbach wanted a longitudinal design, a single journal, a defined cohort, and a systematic attempt to model away the confounders. That's what he built. He took every original research article published in the Proceedings of the National Academy of Sciences between June eighth, 2004 — the day PNAS introduced a one-thousand-dollar option for immediate open access on the journal site — and December twentieth, 2004. That yielded one thousand four hundred ninety-two articles. Of those, two hundred twelve, or fourteen point two percent, were published as immediate open access, meaning the authors paid the fee and the paper was freely available on the PNAS website from day one. The remaining one thousand two hundred eighty, or eighty-five point eight percent, were non-open-access. Same journal, same six-month window, two groups defined by a single decision the authors made about a fee. Eysenbach then addressed the confounders. Open-access articles in this cohort were younger on average — published more recently within the window, at a mean of eighty-three point six days since publication, compared to one hundred four days for non-open-access articles. They had more authors, seven point four on average versus five point seven. They differed in submission track. These aren't trivial differences. More authors mean more co-authors who might cite the work. More recently published papers haven't had as long to accumulate citations. So Eysenbach built regression models that adjusted for all of it: authors' lifetime publication counts and their lifetime average citations per paper, days since publication, number of authors, country of the corresponding author, funding source, subject area, and submission track. Then he measured citations three times — in December 2004, in April 2005, and in October 2005 — to track not just whether open-access articles were cited more, but how quickly the advantage appeared. The findings are striking from the first snapshot. By April fourth, 2005 — a mean of two hundred six days after publication — forty-nine percent of non-open-access articles had received zero citations. For open-access articles, that number was thirty-seven percent. The relative risk for a non-open-access article remaining uncited was one point three, with a ninety-five percent confidence interval of one point one to one point six. That's not dramatic yet. But watch what happens six months later. By October 2005, ten to sixteen months after publication, non-open-access articles were two point six times more likely to still be uncited. Thirteen and a half percent of non-open-access papers had zero citations, compared to just five point two percent of open-access papers. The gap didn't shrink as papers aged into the literature. It grew. Mean citation counts tell the same story. In April 2005, open-access articles averaged one point five citations versus one point two for non-open-access — a modest gap. By October 2005, open-access articles averaged six point four citations while non-open-access articles averaged four point five. That forty-two percent difference in raw counts is a real signal. And when Eysenbach ran the adjusted regression models — controlling for everything he could measure — open-access status remained an independent predictor of being cited. The odds ratio was two point one at the April snapshot, rising to two point nine by October. In plain terms: after accounting for author prominence, team size, discipline, country, and timing, open-access articles were roughly two to three times more likely to have been cited. A survey of two hundred thirty-seven authors — with a seventy-five percent response rate — found no statistically significant differences between open-access and non-open-access authors in their self-rated assessments of their paper's urgency, importance, or quality. That doesn't eliminate selection bias, but it pushes back against the simplest version of it. One of the more telling secondary results is what happens when you break open access into its components. Not all openness is the same. Eysenbach separated papers into four groups: neither open-access nor self-archived, self-archived only, immediately open on the journal site only, and both. By October 2005, papers that were neither self-archived nor immediately open averaged four point four citations. Papers that were self-archived but not immediately open averaged five point two. Papers that were immediately open on the journal site, without self-archiving, averaged six point three. And papers that were both immediately open and self-archived averaged six point five. The pattern is clear: immediate open access on the journal site produced a larger advantage than depositing a copy in PubMed Central or on a personal website. In the regression models, self-archiving lost statistical significance as an independent predictor, while the beta coefficient for immediate journal open access was zero point twenty-five — stronger than the coefficient for self-archiving at zero point fifteen. What this suggests is that discoverability at the point of first encounter matters. If a reader finds a paper and hits a paywall, the friction of locating an archived copy somewhere else is enough to change behavior at the population level. Now, what this study cannot do is prove causation. Eysenbach is clear about that. The design is observational. The regressions adjust for known confounders, but if authors who pay open-access fees differ from other authors in ways not captured by publication count, team size, or discipline — some dimension of ambition or strategic visibility, for example — the study can't rule that out. And the follow-up window is only sixteen months, in a single journal. PNAS is not a typical journal, and Eysenbach acknowledges both points. But here's why the finding still carries weight, even with those caveats. PNAS is one of the most widely subscribed journals in academic science. Its articles become free after six months even without the open-access fee. If you were going to pick a journal where access barriers are already low — where the open-access effect should be hardest to detect — PNAS would be near the top of your list. And Eysenbach found a clear, adjusted, time-increasing advantage anyway. His argument is that if the effect is detectable here, it's likely larger in fields where institutional subscriptions are thinner and the gap between what researchers can access and what they can't is much wider. The policy implication follows directly. Funders and publishers who mandate or fund open access aren't just making a philosophical statement about the ownership of knowledge. They may be changing the actual velocity at which science moves through the world — not by making great papers great, but by ensuring that papers of comparable quality reach more readers faster, and accumulate recognition before the window of peak attention closes. In a scientific culture where citations compound — where being cited early makes you more likely to be cited later — that initial acceleration matters more than it might appear on a single spreadsheet. The quality bias concern hasn't disappeared. Longer follow-up, more journals, and ideally randomized assignment of open-access status would all strengthen the case. But this study gave the question something it had never had before: a controlled, longitudinal, within-journal test. And within those constraints, the advantage was real, it was significant, and it grew over time. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Open access articles get cited more. Not just a little more — in one carefully controlled study, they were nearly three times more likely to be cited by the end of their first year. Now hold that thought for a moment, because the real question isn't the number. It's whether that advantage is real or whether it's just a shadow cast by something else entirely. That something else has a name in the literature: quality bias. The concern is simple and hard to dismiss. Perhaps scientists make their best work freely available because a high-profile result attracts reprint requests, or because a professor puts a landmark paper on a course website, or because it only takes one enthusiastic co-author to upload a preprint. If that's the case, then the association between open access and citations tells you nothing about access at all. It tells you that good papers get read. The question Günther Eysenbach set out to answer is whether open access does anything on its own, stripped of that confounding factor. Every earlier study on this question was cross-sectional — a snapshot comparing citation counts between papers that happened to be freely available and those that weren't, with no ability to control for what made those papers different in the first place. You couldn't tell whether access caused citations or citations caused access. Eysenbach wanted a longitudinal design, a single journal, a defined cohort, and a systematic attempt to model away the confounders. That's what he built.

He took every original research article published in the Proceedings of the National Academy of Sciences between June eighth, 2004 — the day PNAS introduced a one-thousand-dollar option for immediate open access on the journal site — and December twentieth, 2004. That yielded one thousand four hundred ninety-two articles. Of those, two hundred twelve, or fourteen point two percent, were published as immediate open access, meaning the authors paid the fee and the paper was freely available on the PNAS website from day one. The remaining one thousand two hundred eighty, or eighty-five point eight percent, were non-open-access. Same journal, same six-month window, two groups defined by a single decision the authors made about a fee. Eysenbach then addressed the confounders. Open-access articles in this cohort were younger on average — published more recently within the window, at a mean of eighty-three point six days since publication, compared to one hundred four days for non-open-access articles. They had more authors, seven point four on average versus five point seven. They differed in submission track. These aren't trivial differences. More authors mean more co-authors who might cite the work.

More recently published papers haven't had as long to accumulate citations. So Eysenbach built regression models that adjusted for all of it: authors' lifetime publication counts and their lifetime average citations per paper, days since publication, number of authors, country of the corresponding author, funding source, subject area, and submission track. Then he measured citations three times — in December 2004, in April 2005, and in October 2005 — to track not just whether open-access articles were cited more, but how quickly the advantage appeared. The findings are striking from the first snapshot. By April fourth, 2005 — a mean of two hundred six days after publication — forty-nine percent of non-open-access articles had received zero citations. For open-access articles, that number was thirty-seven percent. The relative risk for a non-open-access article remaining uncited was one point three, with a ninety-five percent confidence interval of one point one to one point six. That's not dramatic yet. But watch what happens six months later. By October 2005, ten to sixteen months after publication, non-open-access articles were two point six times more likely to still be uncited. Thirteen and a half percent of non-open-access papers had zero citations, compared to just five point two percent of open-access papers. The gap didn't shrink as papers aged into the literature. It grew.

Mean citation counts tell the same story. In April 2005, open-access articles averaged one point five citations versus one point two for non-open-access — a modest gap. By October 2005, open-access articles averaged six point four citations while non-open-access articles averaged four point five. That forty-two percent difference in raw counts is a real signal. And when Eysenbach ran the adjusted regression models — controlling for everything he could measure — open-access status remained an independent predictor of being cited. The odds ratio was two point one at the April snapshot, rising to two point nine by October. In plain terms: after accounting for author prominence, team size, discipline, country, and timing, open-access articles were roughly two to three times more likely to have been cited. A survey of two hundred thirty-seven authors — with a seventy-five percent response rate — found no statistically significant differences between open-access and non-open-access authors in their self-rated assessments of their paper's urgency, importance, or quality. That doesn't eliminate selection bias, but it pushes back against the simplest version of it. One of the more telling secondary results is what happens when you break open access into its components. Not all openness is the same. Eysenbach separated papers into four groups: neither open-access nor self-archived, self-archived only, immediately open on the journal site only, and both.

By October 2005, papers that were neither self-archived nor immediately open averaged four point four citations. Papers that were self-archived but not immediately open averaged five point two. Papers that were immediately open on the journal site, without self-archiving, averaged six point three. And papers that were both immediately open and self-archived averaged six point five. The pattern is clear: immediate open access on the journal site produced a larger advantage than depositing a copy in PubMed Central or on a personal website. In the regression models, self-archiving lost statistical significance as an independent predictor, while the beta coefficient for immediate journal open access was zero point twenty-five — stronger than the coefficient for self-archiving at zero point fifteen. What this suggests is that discoverability at the point of first encounter matters. If a reader finds a paper and hits a paywall, the friction of locating an archived copy somewhere else is enough to change behavior at the population level. Now, what this study cannot do is prove causation. Eysenbach is clear about that. The design is observational.

The regressions adjust for known confounders, but if authors who pay open-access fees differ from other authors in ways not captured by publication count, team size, or discipline — some dimension of ambition or strategic visibility, for example — the study can't rule that out. And the follow-up window is only sixteen months, in a single journal. PNAS is not a typical journal, and Eysenbach acknowledges both points. But here's why the finding still carries weight, even with those caveats. PNAS is one of the most widely subscribed journals in academic science. Its articles become free after six months even without the open-access fee. If you were going to pick a journal where access barriers are already low — where the open-access effect should be hardest to detect — PNAS would be near the top of your list. And Eysenbach found a clear, adjusted, time-increasing advantage anyway. His argument is that if the effect is detectable here, it's likely larger in fields where institutional subscriptions are thinner and the gap between what researchers can access and what they can't is much wider.

The policy implication follows directly. Funders and publishers who mandate or fund open access aren't just making a philosophical statement about the ownership of knowledge. They may be changing the actual velocity at which science moves through the world — not by making great papers great, but by ensuring that papers of comparable quality reach more readers faster, and accumulate recognition before the window of peak attention closes. In a scientific culture where citations compound — where being cited early makes you more likely to be cited later — that initial acceleration matters more than it might appear on a single spreadsheet. The quality bias concern hasn't disappeared. Longer follow-up, more journals, and ideally randomized assignment of open-access status would all strengthen the case. But this study gave the question something it had never had before: a controlled, longitudinal, within-journal test. And within those constraints, the advantage was real, it was significant, and it grew over time. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Decision Sciences