Reporting Bias in Drug Trials Submitted to the Food and Drug AdministrationReview of Publication and Presentation

Kristin L. Rising, Peter Bacchetti, Lisa BeroView original
OverviewBalancedadam voice
A doctor reads a journal article about a new drug. The study looks clean — it has a favorable primary outcome, a positive conclusion, and has been peer-reviewed. She prescribes it. That chain of trust, from trial to literature to patient, is the backbone of evidence-based medicine. Now, ask a harder question: what if the article she read was a curated version of the real trial? What if the unfavorable outcomes had been quietly dropped, and the conclusion had been rewritten to sound more promising? This is not a hypothetical. A team of researchers measured exactly this by pulling the actual FDA files. Kristin Rising, Peter Bacchetti, and Lisa Bero identified a rare window into the problem. Most research on publication bias runs into a fundamental obstacle — the studies that never appear in journals are, by definition, hard to find. You can compare published trials to registered trials, but registries are incomplete, and the registered record often postdates the trial itself. The innovation here was to use the regulatory record as an independent ground truth. Specifically, the New Drug Application is the dossier a drug company submits to the FDA to obtain permission to market a new drug. By regulation, manufacturers must include the clinical studies on which they rely, so every New Drug Application contains the trials the company considered relevant to demonstrating efficacy. Rising and colleagues pulled all efficacy trials from FDA reviews of New Drug Applications for New Molecular Entities approved from January 2001 through December 2002, then searched for every corresponding publication. That gave them two versions of the evidence: what the FDA saw, and what the medical literature received. They asked two linked questions: how often do the trials submitted to the FDA actually get published? And for the ones that do appear in journals, do the publications faithfully represent what the FDA reviewers saw? Start with publication rates, because the headline number is deceptive. Of the one hundred sixty-four efficacy trials identified in the FDA reviews, one hundred twenty-eight were published — seventy-eight percent. On the surface, that sounds reasonable. Most trials made it into print. But the more important question is which ones didn't, and why. In a multivariate logistic regression, trials with favorable primary outcomes were four point seven times more likely to be published than those without, with an odds ratio of four point seven and a ninety-five percent confidence interval of one point thirty-three to seventeen point one, along with a p-value of zero point zero eighteen. Trials using active comparators were also more likely to appear, with an odds ratio of three point four. It’s worth being precise about what "favorable" means here. A trial was coded as favorable if its primary superiority outcome was statistically significant in favor of the new drug, or if an equivalence outcome was demonstrated. Trials where the drug failed to beat placebo, or where the comparator came out ahead, were coded as not favorable. The four point seven-fold publication advantage for favorable trials means the trials that failed are disproportionately the ones sitting in an FDA filing room, never read by the clinicians who might prescribe the drug. That asymmetry in what gets published is the first filter. The second filter operates inside the papers that do get published, and it’s where things get granular. Rising and colleagues compared the specific outcomes reported in the FDA reviews against those reported in the corresponding published papers, outcome by outcome. Forty-one primary outcomes that appeared in the New Drug Applications were simply omitted from the publications. Prespecified primary endpoints that FDA reviewers saw and evaluated were gone from the journal articles. At the same time, papers were adding outcomes that hadn't been in the regulatory submissions: fifteen additional outcomes appearing only in publications favored the test drug, while only two additional outcomes were neutral or of unknown direction. So the filtering goes both ways — unfavorable outcomes disappear, and favorable outcomes appear. Then, look at what happened to the outcomes that did survive into publication. Excluding outcomes with unknown significance, there were forty-three outcomes in the New Drug Applications that did not favor the drug being reviewed. Of those forty-three, twenty — forty-seven percent — were not included in the published papers. Nearly half of the outcomes where the drug failed simply vanished. Of the remaining twenty-three that appeared in both sources, five had their reported statistical significance change between the regulatory filing and the paper. Four of those five changes went in the direction of favoring the test drug. The drug that didn't beat its comparator in the FDA filing looked like it did in print. Each of these steps is individually small. Forty-one outcomes dropped. Fifteen favorable ones added. Four significance levels quietly shifted. But they compound. The same trial passes through each filter, and at each stage, the result tilts a little further toward the drug looking effective. The finding that lands hardest, though, is at the level of conclusions. Rising and colleagues identified ninety-nine trials where both the FDA review and the published paper provided an explicit conclusion about the drug's efficacy. Nine of those conclusions — nine percent — changed between the regulatory filing and the publication. That might not sound alarming. But every single one of the nine changes went in the direction of favoring the test drug. All nine. Their data show this as one hundred percent, with a ninety-five percent confidence interval of seventy-two to one hundred percent, and a p-value of zero point zero zero thirty-nine. Think about what that means concretely. An FDA reviewer reads a trial and writes: "not superior to comparator." The published paper's conclusion frames the same trial as demonstrating efficacy. The data is the same trial. The regulatory expert saw one thing. The physician reading the journal sees another. And across ninety-nine conclusions, every time there was a discrepancy, it moved in the same direction. This is not random noise. Random noise would produce some changes favoring the drug and some favoring the comparator. What Rising, Bacchetti, and Bero found is a directional lock. Now, put it all together from the doctor's perspective. She never sees the FDA filing. She reads the journal article. The article has dropped nearly half the outcomes where the drug failed. It has added fifteen favorable outcomes that weren't in the regulatory submission. If the drug's primary outcome didn't look good in the New Drug Application, there's a forty-seven percent chance that outcome isn't in the paper she's reading. And if the trial's conclusion changed between the FDA and the journal, it changed to make the drug look better — one hundred percent of the time. Rising and colleagues are careful about the limits of their study. The sample covers two years of approvals, 2001 and 2002, which is a narrow window. Some discrepancies between regulatory submissions and publications could reflect legitimate decisions — condensing a complex trial for a journal audience, for instance. And a seventy-eight percent publication rate is better than some earlier estimates for other contexts. But the directional consistency of every observed distortion is hard to explain as legitimate editorial judgment. Legitimate editing produces noise. This produces a signal. The structural forces driving this are not mysterious. Drug sponsors control whether to submit trial results for publication. Journals, historically, prefer statistically significant positive findings. The FDA record, while technically available under Freedom of Information requests, is cumbersome to access and often presented differently than the original sponsor submission. The researchers who might catch these discrepancies rarely have the time or resources to pull the full New Drug Application and compare it line by line to the published literature. Rising, Bacchetti, and Bero did that work, and the comparison is not reassuring. Their calls for transparency are concrete: full protocols and results posted in trial registries, mandatory reporting of predefined primary outcomes, and result-posting requirements like those introduced by the FDA Amendments Act. The infrastructure for greater transparency existed in nascent form even at the time of this study — ClinicalTrials.gov, the World Health Organization International Clinical Trials Registry Platform, the International Committee of Medical Journal Editors pre-registration requirement. The question was whether those mechanisms would be consistently used and enforced. The twenty-two percent of trials that never appeared in print is one problem. It means the evidence base for approved drugs is already incomplete before you open a single journal. But the distortions inside the seventy-eight percent that did get published compound that incompleteness into something more systematic. The literature that reaches clinicians is not just a partial record. It is a partial record that has been filtered at multiple stages, all in one direction. A doctor reading that literature is not reading the evidence. She is reading the evidence that survived. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

A doctor reads a journal article about a new drug. The study looks clean — it has a favorable primary outcome, a positive conclusion, and has been peer-reviewed. She prescribes it. That chain of trust, from trial to literature to patient, is the backbone of evidence-based medicine. Now, ask a harder question: what if the article she read was a curated version of the real trial? What if the unfavorable outcomes had been quietly dropped, and the conclusion had been rewritten to sound more promising? This is not a hypothetical. A team of researchers measured exactly this by pulling the actual FDA files. Kristin Rising, Peter Bacchetti, and Lisa Bero identified a rare window into the problem. Most research on publication bias runs into a fundamental obstacle — the studies that never appear in journals are, by definition, hard to find. You can compare published trials to registered trials, but registries are incomplete, and the registered record often postdates the trial itself. The innovation here was to use the regulatory record as an independent ground truth. Specifically, the New Drug Application is the dossier a drug company submits to the FDA to obtain permission to market a new drug. By regulation, manufacturers must include the clinical studies on which they rely, so every New Drug Application contains the trials the company considered relevant to demonstrating efficacy.

Rising and colleagues pulled all efficacy trials from FDA reviews of New Drug Applications for New Molecular Entities approved from January 2001 through December 2002, then searched for every corresponding publication. That gave them two versions of the evidence: what the FDA saw, and what the medical literature received. They asked two linked questions: how often do the trials submitted to the FDA actually get published? And for the ones that do appear in journals, do the publications faithfully represent what the FDA reviewers saw? Start with publication rates, because the headline number is deceptive. Of the one hundred sixty-four efficacy trials identified in the FDA reviews, one hundred twenty-eight were published — seventy-eight percent. On the surface, that sounds reasonable. Most trials made it into print. But the more important question is which ones didn't, and why. In a multivariate logistic regression, trials with favorable primary outcomes were four point seven times more likely to be published than those without, with an odds ratio of four point seven and a ninety-five percent confidence interval of one point thirty-three to seventeen point one, along with a p-value of zero point zero eighteen. Trials using active comparators were also more likely to appear, with an odds ratio of three point four. It’s worth being precise about what "favorable" means here.

A trial was coded as favorable if its primary superiority outcome was statistically significant in favor of the new drug, or if an equivalence outcome was demonstrated. Trials where the drug failed to beat placebo, or where the comparator came out ahead, were coded as not favorable. The four point seven-fold publication advantage for favorable trials means the trials that failed are disproportionately the ones sitting in an FDA filing room, never read by the clinicians who might prescribe the drug. That asymmetry in what gets published is the first filter. The second filter operates inside the papers that do get published, and it’s where things get granular. Rising and colleagues compared the specific outcomes reported in the FDA reviews against those reported in the corresponding published papers, outcome by outcome. Forty-one primary outcomes that appeared in the New Drug Applications were simply omitted from the publications. Prespecified primary endpoints that FDA reviewers saw and evaluated were gone from the journal articles. At the same time, papers were adding outcomes that hadn't been in the regulatory submissions: fifteen additional outcomes appearing only in publications favored the test drug, while only two additional outcomes were neutral or of unknown direction. So the filtering goes both ways — unfavorable outcomes disappear, and favorable outcomes appear.

Then, look at what happened to the outcomes that did survive into publication. Excluding outcomes with unknown significance, there were forty-three outcomes in the New Drug Applications that did not favor the drug being reviewed. Of those forty-three, twenty — forty-seven percent — were not included in the published papers. Nearly half of the outcomes where the drug failed simply vanished. Of the remaining twenty-three that appeared in both sources, five had their reported statistical significance change between the regulatory filing and the paper. Four of those five changes went in the direction of favoring the test drug. The drug that didn't beat its comparator in the FDA filing looked like it did in print. Each of these steps is individually small. Forty-one outcomes dropped. Fifteen favorable ones added. Four significance levels quietly shifted. But they compound. The same trial passes through each filter, and at each stage, the result tilts a little further toward the drug looking effective. The finding that lands hardest, though, is at the level of conclusions. Rising and colleagues identified ninety-nine trials where both the FDA review and the published paper provided an explicit conclusion about the drug's efficacy. Nine of those conclusions — nine percent — changed between the regulatory filing and the publication. That might not sound alarming. But every single one of the nine changes went in the direction of favoring the test drug. All nine.

Their data show this as one hundred percent, with a ninety-five percent confidence interval of seventy-two to one hundred percent, and a p-value of zero point zero zero thirty-nine. Think about what that means concretely. An FDA reviewer reads a trial and writes: "not superior to comparator." The published paper's conclusion frames the same trial as demonstrating efficacy. The data is the same trial. The regulatory expert saw one thing. The physician reading the journal sees another. And across ninety-nine conclusions, every time there was a discrepancy, it moved in the same direction. This is not random noise. Random noise would produce some changes favoring the drug and some favoring the comparator. What Rising, Bacchetti, and Bero found is a directional lock. Now, put it all together from the doctor's perspective. She never sees the FDA filing. She reads the journal article. The article has dropped nearly half the outcomes where the drug failed. It has added fifteen favorable outcomes that weren't in the regulatory submission. If the drug's primary outcome didn't look good in the New Drug Application, there's a forty-seven percent chance that outcome isn't in the paper she's reading. And if the trial's conclusion changed between the FDA and the journal, it changed to make the drug look better — one hundred percent of the time.

Rising and colleagues are careful about the limits of their study. The sample covers two years of approvals, 2001 and 2002, which is a narrow window. Some discrepancies between regulatory submissions and publications could reflect legitimate decisions — condensing a complex trial for a journal audience, for instance. And a seventy-eight percent publication rate is better than some earlier estimates for other contexts. But the directional consistency of every observed distortion is hard to explain as legitimate editorial judgment. Legitimate editing produces noise. This produces a signal. The structural forces driving this are not mysterious. Drug sponsors control whether to submit trial results for publication. Journals, historically, prefer statistically significant positive findings. The FDA record, while technically available under Freedom of Information requests, is cumbersome to access and often presented differently than the original sponsor submission. The researchers who might catch these discrepancies rarely have the time or resources to pull the full New Drug Application and compare it line by line to the published literature. Rising, Bacchetti, and Bero did that work, and the comparison is not reassuring.

Their calls for transparency are concrete: full protocols and results posted in trial registries, mandatory reporting of predefined primary outcomes, and result-posting requirements like those introduced by the FDA Amendments Act. The infrastructure for greater transparency existed in nascent form even at the time of this study — ClinicalTrials.gov, the World Health Organization International Clinical Trials Registry Platform, the International Committee of Medical Journal Editors pre-registration requirement. The question was whether those mechanisms would be consistently used and enforced. The twenty-two percent of trials that never appeared in print is one problem. It means the evidence base for approved drugs is already incomplete before you open a single journal. But the distortions inside the seventy-eight percent that did get published compound that incompleteness into something more systematic. The literature that reaches clinicians is not just a partial record. It is a partial record that has been filtered at multiple stages, all in one direction. A doctor reading that literature is not reading the evidence. She is reading the evidence that survived. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Mathematics