Are animal models predictive for humans?

Niall Shanks, Ray Greek, Jean GreekView original
OverviewBalancedalloy voice
When we say an animal model can predict what will happen in humans, what do we really mean? Shanks, Greek, and their colleagues push that question right to the front of the room. Predict isn’t a warm, fuzzy word that means "seems useful" or "lines up with a story we tell after the fact." Predict means you can take the output of a test today and be right about a future human outcome, with high and reliable probability. If you’re going to claim that kind of power for an animal study, they argue, the burden is on you to show it prospectively against human results, not just in anecdotes or correlations that look tidy in hindsight. The stakes here are not abstract. Medicine has a harsh scoreboard because the cost of being wrong is counted in lives and long-term harm. Troglitazone, an antidiabetic, passed through preclinical testing only to cause liver failure in fewer than one percent of users, and it was pulled. Rofecoxib, a painkiller, was withdrawn after heart attacks and strokes were estimated in a similarly small slice of patients. Those numbers are tiny, yet decisive. That’s the level of caution clinicians and regulators live with. So if we’re going to say a model predicts, it has to be measured against the same bar we use anywhere else in medicine. So how do you measure prediction? You don’t do it with general agreement or a gut feel that two species "look similar." You do it with the standard two-by-two framework that pits the animal test against a gold standard, which in this discussion is the human outcome. Picture four boxes: true positives and true negatives, false positives and false negatives. Sensitivity is the share of true human toxicities the animal test catches; that’s true positives divided by true positives plus false negatives. Specificity is the share of safe human outcomes the animal test leaves alone; that’s true negatives divided by true negatives plus false positives. Positive predictive value asks: when the animal test says "toxic," how often is it actually toxic in humans? Negative predictive value asks the mirror question for "safe." Those aren’t just math words. They’re how we tell the difference between something that flags lots of hazards and something that actually helps us decide what will happen to patients. And there’s a trap you want to avoid. Concordance—lots of species showing the same qualitative result—feels reassuring. But high concordance can live alongside poor prediction if the baseline prevalence is low or if there are many false alarms. A noisy smoke detector goes off a lot. That doesn’t mean it tells you which building will burn next. The causal analogical models, or CAMs, perspective pushes this even further. It says that for cross-species prediction to work, the similarities you’re pointing to have to be causally linked to the outcome you care about, and there should be no disanalogies that break that causal chain between the model species and humans. That’s a hard requirement. Genes and alleles, regulation and expression, protein-protein interactions, whole network architectures, organismal physiology, and even the environments those organisms live in—each one is a place where causal wiring can diverge. If the wiring differs in the wrong spot, the same-looking endpoint can sit atop a different mechanism, and your "prediction" starts to wobble. Now, down to the evidence. In toxicity, the signal is mixed at best. In one head-to-head look at six drugs, the animal tests caught about half the human toxicities—sensitivity was zero point fifty-two—and when animals said "toxic," they were right for humans only about a third of the time, with a positive predictive value of zero point thirty-one. That’s closer to coin-flip territory than anyone wants. Older comparisons tell a similar story. In two nineteen nineties datasets, just four of twenty-four toxicities were first found in animals, and in only six of one hundred fourteen clinical toxicities was there an animal correlate that lined up. A Japanese Pharmaceutical Manufacturers Association review of sixty-four marketed drugs found that thirty-nine of ninety-one clinical toxicities were not forecast in animal studies, even when the standard was generous—counting any animal correlate as a "hit." Put differently, if you treat any alignment as success, you still have forty-three percent wrong. That’s not the behavior of a predictor you’d entrust with human safety. What about the big, industry-spanning looks? Olson and a multinational team did what many people point to as the flagship analysis. They pooled data across many compounds and asked: when humans experienced a certain organ toxicity during development, had any preclinical species shown something in that organ? If you frame it that way, you get a true positive concordance of about seventy-one percent when you consider rodents and nonrodents together, about sixty-three percent for nonrodents alone, and about forty-three percent for rodents. Those numbers are often quoted as "predictive." Olson and colleagues were explicit, though: that’s sensitivity-like concordance, not predictive value. Their dataset was limited to compounds that already had human toxicity signals, and the analysis focused on true positives and false negatives. It didn’t populate all four boxes of the predictive table, and without specificity and the rates of false alarms, you can’t calculate the probabilities that matter for clinical decision-making. It tells us something. It does not tell us what a clinician needs to know at the point of care. Carcinogenicity brings the limitation into sharp relief. Using International Agency for Research on Cancer tallies, an older synthesis counted five hundred two agents labeled carcinogenic in animals. Of those, only one hundred four were definite or probable human carcinogens. That’s a positive predictive value of about seven percent. The flip side—three hundred ninety-eight of the five hundred two—were animal carcinogens not deemed definite or probable for humans, a false-positive rate around seventy-nine percent. Later International Agency for Research on Cancer tallies put the fraction of agents classified as definite human carcinogens at about nine point nine percent, and probable at about seven point two percent, for a combined seventeen point one percent as of two thousand four. And there’s a policy echo here. Knight and colleagues compared Environmental Protection Agency and International Agency for Research on Cancer classifications for one hundred twenty-eight chemicals. Where at least limited human data existed, the two lined up; statistically, that alignment wasn’t different from chance with a p-value around zero point fifty-nine. But for one hundred eleven chemicals relying mainly on animal data, the Environmental Protection Agency was far more likely than the International Agency for Research on Cancer to designate higher human risk, with a p-value smaller than one in ten thousand. In other words, leaning on animal data alone tended to over-call human carcinogenicity. Move from hazard to pharmacokinetics, and you meet another source of slippage: bioavailability, the fraction of a dose that reaches systemic circulation. Grass and Sinko, as described in subsequent reviews, plotted bioavailability across humans and several animal species. The picture wasn’t a neat line. It looked like a shotgun blast. Some drugs were highly bioavailable in dogs and not in humans; some the other way around; and similar scatter appeared across primates and rodents. There are patches of correlation, but no single species—or even a tidy combination—gave you a reliably transportable map. The practical meaning is stark. A drug that looks doomed in animals can be viable in humans, and a drug that sails through can still fail when people take it. If your selection gate discards on the basis of animal bioavailability, you risk throwing away future therapies. Mechanism explains a lot of this. Think about phenol metabolism. Humans and rats both clear phenol through sulfate conjugation and glucuronidation, but the balance between those pathways isn’t the same. Cats lean almost entirely on sulfate conjugation because they lack robust glucuronidation; pigs go the other way. Caldwell pointed out that there are metabolic pathways—seven in his count—that are unique to primates. Same endpoint, different routes. That difference matters if a drug or a toxin hits a pathway that’s turned up in one species and down in another. And then there’s thalidomide, the hardest case to ignore. When people realized thalidomide caused characteristic limb defects in human fetuses, a scramble began to see where animals would replicate it. Across roughly ten rat strains, fifteen mouse strains, eleven rabbit breeds, two dog breeds, three hamster strains, eight nonhuman primate species, plus cats, armadillos, guinea pigs, swine, ferrets—you name it—the teratogenic effect could be induced, but only sometimes. In primates, most species except the bushbaby showed the defect, and across fifteen putative human teratogens, eight were teratogenic in one or more primates. Manson and Wise emphasized the kicker: even when you got an animal to mirror the human outcome, you often needed doses twenty-five to one hundred fifty times the human dose, and the usual suspects—absorption, distribution, metabolism, placental transfer—did not cleanly explain the variability. If you tried to design a prospective screen from that patchwork, you would not have had a reliable early warning. Pull the pieces together and you see why some argue that medicine often demands predictive metrics approaching zero point ninety-five to one point zero for sensitivity, specificity, and the positive and negative predictive values. If the consequences of a false negative are catastrophic, or a false positive derails a viable therapy, the tolerance for error collapses. By that yardstick, interspecies tests rarely qualify. And this isn’t just a philosophical worry. As Leavitt reminded readers in two thousand seven, nine out of ten experimental drugs still fail in clinical studies. Preclinical work—animal and otherwise—catches some hazards. It does not give us the clean forecasting tool many people wish we had. The causal analogical models lens helps make sense of that frustration. To be predictive, you need causal sameness where it counts and an absence of the particular disanalogies that scramble cause and effect between species. Genes, regulation, protein interactions, network topology, whole-organism physiology, and environment each represent a place where the wiring can differ, and in complex systems those differences compound. That’s why concordance can be real and still not deliver what a clinician wants: a statement like "this positive means you’ll see the same thing in patients, with high probability," or "this negative really lets you sleep at night." So where does that leave us? First, with language. If a test flags many things animals and humans share, call it concordant. If you want to call it predictive for humans, show sensitivity and specificity, and show the probabilities—positive and negative predictive values—against human outcomes. Olson’s analysis, the International Agency for Research on Cancer tallies, the bioavailability scatterplots, and the case studies from thalidomide to troglitazone all point to the same bottom line: animal studies can surface hazards, sometimes impressively, but as universal predictors of human response they come up short when measured by the standards medicine lives by. Second, with priorities. Shanks, Greek, and colleagues argue that the burden of proof lies with those making predictive claims for animal models, not with skeptics to disprove them. And they suggest leaning into human-relevant, intraspecies approaches—epidemiology, in vitro work with human tissues, carefully designed microdosing studies, and increasingly, in silico models grounded in human data. None of these are magic either. But they attack the problem where the causal wiring matches the target: in humans. That’s the sober reading of the evidence. Animal models can teach us biology, generate hypotheses, and, in some domains, catch dangers we might otherwise miss. What they have not done, consistently and with high probability, is forecast human outcomes in the way clinicians, regulators, and patients mean when they say the word predict. If we reserve that word for tools that clear the predictive bar, and build the rest of our toolkit around human-grounded methods and mechanistic understanding, we get both honesty about where we are and a path to being less wrong tomorrow.

When we say an animal model can predict what will happen in humans, what do we really mean? Shanks, Greek, and their colleagues push that question right to the front of the room. Predict isn’t a warm, fuzzy word that means "seems useful" or "lines up with a story we tell after the fact." Predict means you can take the output of a test today and be right about a future human outcome, with high and reliable probability.

If you’re going to claim that kind of power for an animal study, they argue, the burden is on you to show it prospectively against human results, not just in anecdotes or correlations that look tidy in hindsight.

The stakes here are not abstract. Medicine has a harsh scoreboard because the cost of being wrong is counted in lives and long-term harm. Troglitazone, an antidiabetic, passed through preclinical testing only to cause liver failure in fewer than one percent of users, and it was pulled.

Rofecoxib, a painkiller, was withdrawn after heart attacks and strokes were estimated in a similarly small slice of patients. Those numbers are tiny, yet decisive. That’s the level of caution clinicians and regulators live with.

So if we’re going to say a model predicts, it has to be measured against the same bar we use anywhere else in medicine.

So how do you measure prediction? You don’t do it with general agreement or a gut feel that two species "look similar." You do it with the standard two-by-two framework that pits the animal test against a gold standard, which in this discussion is the human outcome. Picture four boxes: true positives and true negatives, false positives and false negatives.

Sensitivity is the share of true human toxicities the animal test catches; that’s true positives divided by true positives plus false negatives. Specificity is the share of safe human outcomes the animal test leaves alone; that’s true negatives divided by true negatives plus false positives. Positive predictive value asks: when the animal test says "toxic," how often is it actually toxic in humans?

Negative predictive value asks the mirror question for "safe." Those aren’t just math words. They’re how we tell the difference between something that flags lots of hazards and something that actually helps us decide what will happen to patients.

And there’s a trap you want to avoid. Concordance—lots of species showing the same qualitative result—feels reassuring. But high concordance can live alongside poor prediction if the baseline prevalence is low or if there are many false alarms.

A noisy smoke detector goes off a lot. That doesn’t mean it tells you which building will burn next.

The causal analogical models, or CAMs, perspective pushes this even further. It says that for cross-species prediction to work, the similarities you’re pointing to have to be causally linked to the outcome you care about, and there should be no disanalogies that break that causal chain between the model species and humans. That’s a hard requirement.

Genes and alleles, regulation and expression, protein-protein interactions, whole network architectures, organismal physiology, and even the environments those organisms live in—each one is a place where causal wiring can diverge. If the wiring differs in the wrong spot, the same-looking endpoint can sit atop a different mechanism, and your "prediction" starts to wobble.

Now, down to the evidence. In toxicity, the signal is mixed at best. In one head-to-head look at six drugs, the animal tests caught about half the human toxicities—sensitivity was zero point fifty-two—and when animals said "toxic," they were right for humans only about a third of the time, with a positive predictive value of zero point thirty-one.

That’s closer to coin-flip territory than anyone wants. Older comparisons tell a similar story. In two nineteen nineties datasets, just four of twenty-four toxicities were first found in animals, and in only six of one hundred fourteen clinical toxicities was there an animal correlate that lined up.

A Japanese Pharmaceutical Manufacturers Association review of sixty-four marketed drugs found that thirty-nine of ninety-one clinical toxicities were not forecast in animal studies, even when the standard was generous—counting any animal correlate as a "hit." Put differently, if you treat any alignment as success, you still have forty-three percent wrong. That’s not the behavior of a predictor you’d entrust with human safety.

What about the big, industry-spanning looks? Olson and a multinational team did what many people point to as the flagship analysis. They pooled data across many compounds and asked: when humans experienced a certain organ toxicity during development, had any preclinical species shown something in that organ?

If you frame it that way, you get a true positive concordance of about seventy-one percent when you consider rodents and nonrodents together, about sixty-three percent for nonrodents alone, and about forty-three percent for rodents. Those numbers are often quoted as "predictive." Olson and colleagues were explicit, though: that’s sensitivity-like concordance, not predictive value. Their dataset was limited to compounds that already had human toxicity signals, and the analysis focused on true positives and false negatives.

It didn’t populate all four boxes of the predictive table, and without specificity and the rates of false alarms, you can’t calculate the probabilities that matter for clinical decision-making. It tells us something. It does not tell us what a clinician needs to know at the point of care.

Carcinogenicity brings the limitation into sharp relief. Using International Agency for Research on Cancer tallies, an older synthesis counted five hundred two agents labeled carcinogenic in animals. Of those, only one hundred four were definite or probable human carcinogens.

That’s a positive predictive value of about seven percent. The flip side—three hundred ninety-eight of the five hundred two—were animal carcinogens not deemed definite or probable for humans, a false-positive rate around seventy-nine percent. Later International Agency for Research on Cancer tallies put the fraction of agents classified as definite human carcinogens at about nine point nine percent, and probable at about seven point two percent, for a combined seventeen point one percent as of two thousand four.

And there’s a policy echo here. Knight and colleagues compared Environmental Protection Agency and International Agency for Research on Cancer classifications for one hundred twenty-eight chemicals. Where at least limited human data existed, the two lined up; statistically, that alignment wasn’t different from chance with a p-value around zero point fifty-nine.

But for one hundred eleven chemicals relying mainly on animal data, the Environmental Protection Agency was far more likely than the International Agency for Research on Cancer to designate higher human risk, with a p-value smaller than one in ten thousand. In other words, leaning on animal data alone tended to over-call human carcinogenicity.

Move from hazard to pharmacokinetics, and you meet another source of slippage: bioavailability, the fraction of a dose that reaches systemic circulation. Grass and Sinko, as described in subsequent reviews, plotted bioavailability across humans and several animal species. The picture wasn’t a neat line.

It looked like a shotgun blast. Some drugs were highly bioavailable in dogs and not in humans; some the other way around; and similar scatter appeared across primates and rodents. There are patches of correlation, but no single species—or even a tidy combination—gave you a reliably transportable map.

The practical meaning is stark. A drug that looks doomed in animals can be viable in humans, and a drug that sails through can still fail when people take it. If your selection gate discards on the basis of animal bioavailability, you risk throwing away future therapies.

Mechanism explains a lot of this. Think about phenol metabolism. Humans and rats both clear phenol through sulfate conjugation and glucuronidation, but the balance between those pathways isn’t the same.

Cats lean almost entirely on sulfate conjugation because they lack robust glucuronidation; pigs go the other way. Caldwell pointed out that there are metabolic pathways—seven in his count—that are unique to primates. Same endpoint, different routes.

That difference matters if a drug or a toxin hits a pathway that’s turned up in one species and down in another.

And then there’s thalidomide, the hardest case to ignore. When people realized thalidomide caused characteristic limb defects in human fetuses, a scramble began to see where animals would replicate it. Across roughly ten rat strains, fifteen mouse strains, eleven rabbit breeds, two dog breeds, three hamster strains, eight nonhuman primate species, plus cats, armadillos, guinea pigs, swine, ferrets—you name it—the teratogenic effect could be induced, but only sometimes.

In primates, most species except the bushbaby showed the defect, and across fifteen putative human teratogens, eight were teratogenic in one or more primates. Manson and Wise emphasized the kicker: even when you got an animal to mirror the human outcome, you often needed doses twenty-five to one hundred fifty times the human dose, and the usual suspects—absorption, distribution, metabolism, placental transfer—did not cleanly explain the variability. If you tried to design a prospective screen from that patchwork, you would not have had a reliable early warning.

Pull the pieces together and you see why some argue that medicine often demands predictive metrics approaching zero point ninety-five to one point zero for sensitivity, specificity, and the positive and negative predictive values. If the consequences of a false negative are catastrophic, or a false positive derails a viable therapy, the tolerance for error collapses. By that yardstick, interspecies tests rarely qualify.

And this isn’t just a philosophical worry. As Leavitt reminded readers in two thousand seven, nine out of ten experimental drugs still fail in clinical studies. Preclinical work—animal and otherwise—catches some hazards. It does not give us the clean forecasting tool many people wish we had.

The causal analogical models lens helps make sense of that frustration. To be predictive, you need causal sameness where it counts and an absence of the particular disanalogies that scramble cause and effect between species. Genes, regulation, protein interactions, network topology, whole-organism physiology, and environment each represent a place where the wiring can differ, and in complex systems those differences compound.

That’s why concordance can be real and still not deliver what a clinician wants: a statement like "this positive means you’ll see the same thing in patients, with high probability," or "this negative really lets you sleep at night."

So where does that leave us? First, with language. If a test flags many things animals and humans share, call it concordant.

If you want to call it predictive for humans, show sensitivity and specificity, and show the probabilities—positive and negative predictive values—against human outcomes. Olson’s analysis, the International Agency for Research on Cancer tallies, the bioavailability scatterplots, and the case studies from thalidomide to troglitazone all point to the same bottom line: animal studies can surface hazards, sometimes impressively, but as universal predictors of human response they come up short when measured by the standards medicine lives by.

Second, with priorities. Shanks, Greek, and colleagues argue that the burden of proof lies with those making predictive claims for animal models, not with skeptics to disprove them. And they suggest leaning into human-relevant, intraspecies approaches—epidemiology, in vitro work with human tissues, carefully designed microdosing studies, and increasingly, in silico models grounded in human data.

None of these are magic either. But they attack the problem where the causal wiring matches the target: in humans.

That’s the sober reading of the evidence. Animal models can teach us biology, generate hypotheses, and, in some domains, catch dangers we might otherwise miss. What they have not done, consistently and with high probability, is forecast human outcomes in the way clinicians, regulators, and patients mean when they say the word predict.

If we reserve that word for tools that clear the predictive bar, and build the rest of our toolkit around human-grounded methods and mechanistic understanding, we get both honesty about where we are and a path to being less wrong tomorrow.

More in Veterinary