How Many Scientists Fabricate and Falsify Research? A Systematic Review and Meta-Analysis of Survey Data

Daniele FanelliView original
OverviewBalancedalloy voice
Imagine you're walking into a lab where the shelves are lined with decades of experiments. You're trusting that what's in those bottles and spreadsheets reflects what really happened. That trust is the quiet engine of science. So when we talk about misconduct—people making up data, changing results, or massaging analyses until they look good—we're not just gossiping about bad apples. We're asking whether the library we all rely on has missing pages or forged lines. For years, the conversation swung between shocking anecdotes and hand waving reassurance. What was missing was a careful, quantitative picture. Daniele Fanelli stepped into that gap in 2009 with something we hadn't had before: a systematic review and meta-analysis that tried to harmonize a messy stack of surveys into numbers you could actually compare. He drew a clean boundary around the behaviors that directly distort knowledge. Fabrication is inventing data or cases out of thin air. Falsification is distorting real data or results. And "modification"—sometimes called cooking the data—covers practices that alter findings to make outcomes look better, like selectively reporting or mining for p-values. There's a continuum here with all kinds of questionable research practices, but for the main analysis, Fanelli kept the focus on these serious distortions. Things like plagiarism, gift authorship, or abusive supervision were set aside. Different problem, different toolkit. If you've ever tried to compare surveys, you know how slippery this can be. One study asks "Have you ever falsified data?" while another says "Have you ever modified or selectively reported findings?" One is mailed, another is handed out in a room. Some ask about you; others ask what you know about your colleagues. Fanelli's move was to standardize the outcome across this jungle: for each question, turn it into the proportion of people who recalled at least one incident—either an admission about themselves or knowledge of a colleague doing it. To make that work, he only included surveys that had a clear "never" or "none" category, so there was a clean baseline. Undergraduate misconduct was excluded—too far from the realm of publishable science—and only quantitative frequencies made the cut. The study base ended up covering two decades of work. Twenty-one surveys passed screening, with eighteen entering the main meta-analysis. They spanned from nineteen eighty-seven through two thousand eight, and most respondents were in the United States, roughly seven in ten. That mix matters because any pooled estimate is always a product of what you feed into it—discipline, country, era, and how the question was put. Statistically, the plan was straightforward and transparent. Fanelli used random-effects models to pool estimates, because the studies weren't clones of each other; they were different windows on the same landscape. Heterogeneity—the fact that the numbers varied more than you'd expect from chance—was not a footnote; it was the headline. A test called Cochran's Q flagged substantial between-study variation. Rather than shrug, Fanelli turned to meta-regression, asking which methodological choices were driving those differences. Three stood out: whether you asked about yourself or your colleagues, whether the survey was mailed or handed out, and whether the wording explicitly used the words "fabrication" or "falsification" versus broader terms like "modification." Together, those factors explained a large share of the variance across studies—on the order of eighty percent—and the relationships were statistically strong. Sensitivity checks that dropped one study at a time told the same story, and when a particularly influential paper by Titus and colleagues was removed, the model's fit improved further, with explained variance rising to about ninety-two percent. Publication bias is always a worry when you don't have many studies, and they don't look alike, so Fanelli treated funnel plots with caution and emphasized interpretation over overconfidence. So what did the numbers say when you put it all together? Start with the hard core: self-admitted fabrication, falsification, or alteration of data. Across the included surveys, about one point ninety-seven percent of researchers said they had done this at least once, with a confidence interval running roughly from under one percent to about four point five percent. That's small—and it should be. But it is not zero. When the questions spelled out "fabrication" or "falsification" explicitly, the pooled estimate dropped to about one point zero six percent. Tighter wording depressed admissions, but didn't make them vanish. Step back to the broader gray zone—the questionable research practices that might not be outright fakery but can nudge results in misleading ways. In that space, self-admissions rose sharply. Across a set of questions in six studies, the crude average was nine point fifty-four percent, and in some items up to a third of respondents said they had done at least one such thing. That is a lot of room for bias to creep into the literature, even before you get to the dramatic cases. Now turn the camera outward. When the same kind of questions asked about colleagues instead of the self, the numbers jumped. Knowledge of a peer engaging in fabrication, falsification, or alteration pooled at fourteen point twelve percent, and when the wording was restricted to the explicit "fabrication or falsification" terms, it settled near twelve point thirty-four percent. That gap between roughly two percent for self-admissions and the low teens for colleagues isn't surprising—people tend to underreport about themselves and overreport about others—but it's telling. It says there's a signal that doesn't depend on a single survey or a single field. And when you look at knowledge of broader questionable practices among colleagues, the ceiling rises further, with some items approaching three-quarters of respondents saying they knew of such behavior. If you worry that any single study might be tilting the averages, the sensitivity analysis helps. Remove one study at a time, and the self-admission rate for serious misconduct ranged from about one point sixty-five to two point ninety-three percent. The colleague reports barely budged, landing between about twelve point eighty-five and fifteen point forty-one percent. Those aren't wild swings; they're wiggles around a stable signal. There's another way to validate the picture, and that's to change the statistical scale. On a non-logit scale—closer to the everyday percent you and I think in—Fanelli's pooled estimate for self-admitted fabrication, falsification, or alteration was two point thirty-three percent, and the colleague figure was fourteen point forty-eight percent. And if you push the design to be maximally conservative—only mailed surveys, only self-reports, and only questions that name "fabrication" or "falsification"—you can get the number down to zero point sixty-four percent. That's the floor that the strictest methods produce. Floors are useful; they keep us from falling. But they're not where most people live. One of the clearest takeaways is that survey design is not a neutral backdrop; it's an active ingredient. Self-report depresses admissions compared to asking about colleagues. Mailing a questionnaire tends to produce lower rates than handing one out. Using explicit terms like "fabrication" and "falsification" yields lower numbers than broader phrases like "modification." In Fanelli's regression, these factors were all statistically significant drivers of variation, and together they explained most of why study A got a different number than study B. That tells you two things at once: first, you can't read a standalone headline figure and assume it maps to some deep law of human behavior; second, if you standardize the way you ask, you get a clearer view of the underlying phenomenon. Field differences add another layer. After controlling for those methodological factors, samples drawn from medical, clinical, or pharmacological research tended to report higher rates. That effect wasn't bulletproof across every sensitivity check, but it showed up enough to be noticed. And remember that influential Titus paper? Its leverage on the model reminded everyone that single large or distinctive studies can sway meta-analytic portraits, which is why Fanelli kept rerunning the models to check stability. Every study like this sits in a fog of caveats, and Fanelli is clear about them. Non-response bias is the perennial ghost: if people who have something to hide are less likely to answer the survey, you'll underestimate misconduct. Social desirability plays in the opposite direction—people might shade their answers to look virtuous. The boundary between outright falsification and "just" sloppy or opportunistic analysis can be fuzzy in a respondent's mind, especially across disciplines with different norms. And with a small number of heterogeneous studies, classic tools for spotting publication bias don't have much power. There was even one survey that used a randomized response technique—an anonymizing method designed to reduce self-censorship—and in that handed-out format, they didn't find evidence of non-response bias. That's comforting for that specific setup, but it doesn't solve the broader problem. So where does this leave us? With a conservative baseline, and a set of dials we now know matter. On the numbers, self-admitted serious misconduct lands around two percent in pooled analyses, and you can push it down to well under one percent with the strictest survey design. Knowledge of colleagues' serious misconduct sits an order of magnitude higher, in the low teens, and expands sharply when you include the broader questionable practices that tilt the playing field before a single data point is fabricated. On the methods, the three big dials—self versus colleague reporting, mailed versus handed delivery, and explicit versus broad wording—explain most of the disagreement across surveys. Add in time period, country, and field, and you get the textured map Fanelli drew in 2009. If you care about protecting the scientific record, that map is the workhorse. It tells journal editors and funders that better measurement isn't a luxury; it's the only way to know whether interventions are working. It tells researchers designing the next wave of surveys to standardize wording, maximize anonymity, and report enough detail that their results can be compared and pooled. And it tells the rest of us—readers, reviewers, the scientifically curious—that a small, non-zero rate of serious misconduct coexists with a much larger universe of corner cutting that can still warp the published record. A final thought, and then we'll close. It's tempting to jump from these numbers to grand pronouncements about the soul of science. Don't. Fanelli's synthesis doesn't excuse or justify anything; it gives us a baseline. The right next moves are technical and unglamorous: tighten survey instruments, expand coverage beyond the United States, and build field-specific monitoring where the risks and incentives differ. If we do that, the next time we look at this shelf of experiments, we'll be looking with clearer eyes—and better tools to keep it honest.

Imagine you're walking into a lab where the shelves are lined with decades of experiments. You're trusting that what's in those bottles and spreadsheets reflects what really happened. That trust is the quiet engine of science.

So when we talk about misconduct—people making up data, changing results, or massaging analyses until they look good—we're not just gossiping about bad apples. We're asking whether the library we all rely on has missing pages or forged lines. For years, the conversation swung between shocking anecdotes and hand waving reassurance. What was missing was a careful, quantitative picture.

Daniele Fanelli stepped into that gap in 2009 with something we hadn't had before: a systematic review and meta-analysis that tried to harmonize a messy stack of surveys into numbers you could actually compare. He drew a clean boundary around the behaviors that directly distort knowledge. Fabrication is inventing data or cases out of thin air.

Falsification is distorting real data or results. And "modification"—sometimes called cooking the data—covers practices that alter findings to make outcomes look better, like selectively reporting or mining for p-values. There's a continuum here with all kinds of questionable research practices, but for the main analysis, Fanelli kept the focus on these serious distortions.

Things like plagiarism, gift authorship, or abusive supervision were set aside. Different problem, different toolkit.

If you've ever tried to compare surveys, you know how slippery this can be. One study asks "Have you ever falsified data?" while another says "Have you ever modified or selectively reported findings?" One is mailed, another is handed out in a room. Some ask about you; others ask what you know about your colleagues.

Fanelli's move was to standardize the outcome across this jungle: for each question, turn it into the proportion of people who recalled at least one incident—either an admission about themselves or knowledge of a colleague doing it. To make that work, he only included surveys that had a clear "never" or "none" category, so there was a clean baseline. Undergraduate misconduct was excluded—too far from the realm of publishable science—and only quantitative frequencies made the cut.

The study base ended up covering two decades of work. Twenty-one surveys passed screening, with eighteen entering the main meta-analysis. They spanned from nineteen eighty-seven through two thousand eight, and most respondents were in the United States, roughly seven in ten.

That mix matters because any pooled estimate is always a product of what you feed into it—discipline, country, era, and how the question was put.

Statistically, the plan was straightforward and transparent. Fanelli used random-effects models to pool estimates, because the studies weren't clones of each other; they were different windows on the same landscape. Heterogeneity—the fact that the numbers varied more than you'd expect from chance—was not a footnote; it was the headline.

A test called Cochran's Q flagged substantial between-study variation. Rather than shrug, Fanelli turned to meta-regression, asking which methodological choices were driving those differences. Three stood out: whether you asked about yourself or your colleagues, whether the survey was mailed or handed out, and whether the wording explicitly used the words "fabrication" or "falsification" versus broader terms like "modification." Together, those factors explained a large share of the variance across studies—on the order of eighty percent—and the relationships were statistically strong.

Sensitivity checks that dropped one study at a time told the same story, and when a particularly influential paper by Titus and colleagues was removed, the model's fit improved further, with explained variance rising to about ninety-two percent. Publication bias is always a worry when you don't have many studies, and they don't look alike, so Fanelli treated funnel plots with caution and emphasized interpretation over overconfidence.

So what did the numbers say when you put it all together? Start with the hard core: self-admitted fabrication, falsification, or alteration of data. Across the included surveys, about one point ninety-seven percent of researchers said they had done this at least once, with a confidence interval running roughly from under one percent to about four point five percent.

That's small—and it should be. But it is not zero. When the questions spelled out "fabrication" or "falsification" explicitly, the pooled estimate dropped to about one point zero six percent. Tighter wording depressed admissions, but didn't make them vanish.

Step back to the broader gray zone—the questionable research practices that might not be outright fakery but can nudge results in misleading ways. In that space, self-admissions rose sharply. Across a set of questions in six studies, the crude average was nine point fifty-four percent, and in some items up to a third of respondents said they had done at least one such thing.

That is a lot of room for bias to creep into the literature, even before you get to the dramatic cases.

Now turn the camera outward. When the same kind of questions asked about colleagues instead of the self, the numbers jumped. Knowledge of a peer engaging in fabrication, falsification, or alteration pooled at fourteen point twelve percent, and when the wording was restricted to the explicit "fabrication or falsification" terms, it settled near twelve point thirty-four percent.

That gap between roughly two percent for self-admissions and the low teens for colleagues isn't surprising—people tend to underreport about themselves and overreport about others—but it's telling. It says there's a signal that doesn't depend on a single survey or a single field. And when you look at knowledge of broader questionable practices among colleagues, the ceiling rises further, with some items approaching three-quarters of respondents saying they knew of such behavior.

If you worry that any single study might be tilting the averages, the sensitivity analysis helps. Remove one study at a time, and the self-admission rate for serious misconduct ranged from about one point sixty-five to two point ninety-three percent. The colleague reports barely budged, landing between about twelve point eighty-five and fifteen point forty-one percent. Those aren't wild swings; they're wiggles around a stable signal.

There's another way to validate the picture, and that's to change the statistical scale. On a non-logit scale—closer to the everyday percent you and I think in—Fanelli's pooled estimate for self-admitted fabrication, falsification, or alteration was two point thirty-three percent, and the colleague figure was fourteen point forty-eight percent. And if you push the design to be maximally conservative—only mailed surveys, only self-reports, and only questions that name "fabrication" or "falsification"—you can get the number down to zero point sixty-four percent.

That's the floor that the strictest methods produce. Floors are useful; they keep us from falling. But they're not where most people live.

One of the clearest takeaways is that survey design is not a neutral backdrop; it's an active ingredient. Self-report depresses admissions compared to asking about colleagues. Mailing a questionnaire tends to produce lower rates than handing one out.

Using explicit terms like "fabrication" and "falsification" yields lower numbers than broader phrases like "modification." In Fanelli's regression, these factors were all statistically significant drivers of variation, and together they explained most of why study A got a different number than study B. That tells you two things at once: first, you can't read a standalone headline figure and assume it maps to some deep law of human behavior; second, if you standardize the way you ask, you get a clearer view of the underlying phenomenon.

Field differences add another layer. After controlling for those methodological factors, samples drawn from medical, clinical, or pharmacological research tended to report higher rates. That effect wasn't bulletproof across every sensitivity check, but it showed up enough to be noticed.

And remember that influential Titus paper? Its leverage on the model reminded everyone that single large or distinctive studies can sway meta-analytic portraits, which is why Fanelli kept rerunning the models to check stability.

Every study like this sits in a fog of caveats, and Fanelli is clear about them. Non-response bias is the perennial ghost: if people who have something to hide are less likely to answer the survey, you'll underestimate misconduct. Social desirability plays in the opposite direction—people might shade their answers to look virtuous.

The boundary between outright falsification and "just" sloppy or opportunistic analysis can be fuzzy in a respondent's mind, especially across disciplines with different norms. And with a small number of heterogeneous studies, classic tools for spotting publication bias don't have much power. There was even one survey that used a randomized response technique—an anonymizing method designed to reduce self-censorship—and in that handed-out format, they didn't find evidence of non-response bias.

That's comforting for that specific setup, but it doesn't solve the broader problem.

So where does this leave us? With a conservative baseline, and a set of dials we now know matter. On the numbers, self-admitted serious misconduct lands around two percent in pooled analyses, and you can push it down to well under one percent with the strictest survey design.

Knowledge of colleagues' serious misconduct sits an order of magnitude higher, in the low teens, and expands sharply when you include the broader questionable practices that tilt the playing field before a single data point is fabricated. On the methods, the three big dials—self versus colleague reporting, mailed versus handed delivery, and explicit versus broad wording—explain most of the disagreement across surveys. Add in time period, country, and field, and you get the textured map Fanelli drew in 2009.

If you care about protecting the scientific record, that map is the workhorse. It tells journal editors and funders that better measurement isn't a luxury; it's the only way to know whether interventions are working. It tells researchers designing the next wave of surveys to standardize wording, maximize anonymity, and report enough detail that their results can be compared and pooled.

And it tells the rest of us—readers, reviewers, the scientifically curious—that a small, non-zero rate of serious misconduct coexists with a much larger universe of corner cutting that can still warp the published record.

A final thought, and then we'll close. It's tempting to jump from these numbers to grand pronouncements about the soul of science. Don't.

Fanelli's synthesis doesn't excuse or justify anything; it gives us a baseline. The right next moves are technical and unglamorous: tighten survey instruments, expand coverage beyond the United States, and build field-specific monitoring where the risks and incentives differ. If we do that, the next time we look at this shelf of experiments, we'll be looking with clearer eyes—and better tools to keep it honest.

More in Social Sciences