Statistically Controlling for Confounding Constructs Is Harder than You Think
Imagine a researcher who has spent three years developing a new personality questionnaire. They run a study, collect data from several hundred participants, enter their new measure and a competing one into a regression, and find that their measure still predicts job performance even after statistically controlling for the competitor. They write it up, it gets published, and the finding enters the literature as established fact. According to Jacob Westfall and Tal Yarkoni, there is a good chance that finding was never real — not because of fraud or carelessness, but because the statistical procedure everyone uses for this kind of claim is broken in a way that gets worse, not better, as you collect more data. That is the core argument of their paper, "Statistically Controlling for Confounding Constructs Is Harder than You Think." And it has implications for a large fraction of findings across social and behavioral science. The logic of incremental validity is simple and appealing. You have two predictors and an outcome. You want to show that your predictor contributes something unique, above and beyond what the competing predictor already explains. So you run a multiple regression, include the competitor as a covariate, and check whether your variable still has a nonzero coefficient. If it does, you conclude your construct adds something independent. Westfall and Yarkoni note that literally hundreds of thousands of studies have relied on exactly this procedure.
It is the backbone of applied personality research, clinical assessment, educational psychology, and much else. The procedure feels mathematically rigorous. That is precisely what makes the problem so serious. The hidden flaw is measurement error, and here is why it matters. Psychological constructs are never observed directly. You observe a noisy proxy — a questionnaire, a rating scale, a single survey item. That proxy is only a blurry version of the underlying latent construct. When you put the noisy proxy into a regression as a covariate, you remove only the variance captured by that blurry version. You do not remove all the variance associated with the true underlying construct. A residue of the true confound remains, and that residue gets misattributed to your predictor of interest. Westfall and Yarkoni make this concrete with a deliberately absurd but revealing example. Suppose the true confound is daily temperature, but all you have is a single self-reported Likert rating of perceived heat, with a reliability of 0.40 — typical for a single item in psychology. After controlling for those noisy ratings, the correlation between ice cream sales and swimming pool deaths remained positive and statistically significant.
After controlling for the actual recorded daily temperatures, the relationship vanished entirely. Same data, same outcome, different covariate quality — completely different conclusion. Controlling for a noisy proxy is not the same as controlling for the construct. The regression cannot remove what it cannot see. This intuition becomes a formal framework in their paper, one that links reliability, correlations between predictors and outcomes, and sample size to determine how often the standard regression test will produce false positives. And the numbers that emerge from that framework are alarming. Westfall and Yarkoni ran Monte Carlo simulations — computer experiments generating fake data under a known null hypothesis, then checking how often the regression incorrectly rejected it. Their headline finding is counterintuitive: Type I error rates, the rate of false positives, are highest when sample sizes are large and measurement reliability is only moderate. Larger samples normally mean better science. Here, larger samples mean false positives pile up faster, because increased precision sharpens the biased estimate rather than correcting it.
A few numbers anchor the pattern. Repeating the ice cream simulation ten thousand times, Westfall and Yarkoni found a spuriously significant result ninety-two percent of the time. In their broader parameter table, with a sample size of three thousand participants, predictor-covariate correlations of 0.3, and reliability of 0.40, the Type I error rate was forty-eight percent. Push the sample to thirty thousand participants and the error rate climbs to seventy-six percent. By contrast, at a small sample of thirty participants with the same reliability, the error rate was twelve percent — still above the nominal five percent, but nowhere near the catastrophic rates at larger sample sizes. Higher reliability helps substantially: at thirty thousand participants and reliability of 0.80, the error rate drops to twenty-five percent. Methods that model measurement error directly can maintain the nominal five percent. The pattern is consistent and stark: noisy covariates plus big samples are a recipe for mass false positives.
Westfall and Yarkoni then moved from simulation to a real dataset to show the problem is not merely theoretical. They used the Eugene-Springfield community sample — six hundred and four participants who completed both the NEO Personality Inventory Revised, the standard Big Five measure, and the HEXACO Personality Inventory, along with a Behavioral Report Inventory covering roughly four hundred specific activities collapsed into sixty behavior clusters. The key question was whether the HEXACO model's Honesty-Humility factor predicts behavior above and beyond the five NEO factors. They compared what ordinary multiple regression said to what a structural equation model — or SEM — said. Structural equation modeling explicitly represents latent constructs as separate from their measured indicators. Rather than entering observed scale scores directly, an SEM estimates the true-score variance for each construct and models how the measured items relate to that latent signal. The control in an SEM is a control for the actual latent construct, not a blurry proxy. The two approaches told different stories. Standard regression suggested that HEXACO and NEO versions of several personality factors made independent contributions to behavior — they looked like separable constructs. The structural equation model substantially attenuated or eliminated those separable effects for most factors, consistent with what you would expect when measurement error is properly partitioned out.
For the focal Honesty-Humility question, both approaches identified some behaviors where Honesty-Humility appeared to predict outcomes beyond the NEO factors. But Westfall and Yarkoni show this agreement is fragile. When they reran the structural equation model treating each personality factor as a single measured indicator and varied the assumed reliability, the significance of Honesty-Humility's incremental effect depended heavily on that assumption. It only held when assumed reliabilities were at least 0.78. For comparison, Neuroticism's effect held as long as reliabilities were at least 0.50. The observed reliabilities in the dataset ranged from 0.66 to 0.79, with a mean around 0.74 — right at the edge of where Honesty-Humility's incremental validity would register at all. The empirical finding is genuinely fragile. So, structural equation modeling is the right tool. But it comes with real costs, and Westfall and Yarkoni are explicit about them. Structural equation modeling requires either multiple measured indicators per construct — so reliabilities can be estimated from the data — or explicit, justified reliability assumptions for any single-indicator measure.
When you have only one item measuring a construct, like parental education assessed with a single survey question, you must fix the measurement error in your model based on an assumed reliability, and you must report sensitivity analyses over a plausible range of those assumptions. Your conclusions can hinge on them entirely. The power requirements are also much steeper than most researchers expect. For eighty percent power to detect a partial effect of moderate size, you need roughly two hundred participants when the covariate is perfectly reliable. If reliability drops to 0.40, that requirement balloons to around two thousand three hundred participants for the same effect. Detecting a small unique contribution may require tens of thousands. These are not sample sizes that most psychological studies even approach. Westfall and Yarkoni built a web application that lets researchers explore these tradeoffs for their own parameter values — a practical tool for anyone trying to plan or evaluate an incremental validity study. The broader implication is uncomfortable but important. A potentially large proportion of published incremental validity claims were made using regression on noisy measures, in samples large enough to make false positives nearly certain, without any correction for measurement error. Those findings are in the literature, cited, built upon.
The fix is available: use structural equation modeling, collect multiple indicators, make reliability assumptions explicit. But the fix is harder and more demanding than the procedure it replaces. That is precisely why it has been ignored. When you read a paper claiming that one construct predicts an outcome "above and beyond" another, ask two questions: Did the authors account for measurement error? Did they use a latent-variable approach? If the answer to both is no — and in most published work, it is — the incremental validity claim is much weaker than it appears. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- The Cognitive Structure of Emotions.
- The dark side of creativity: Original thinkers can be more dishonest.
- Niveles de estrés, ansiedad y depresión en la primera fase del brote del COVID-19 en una muestra recogida en el norte de España
- Understanding Learning Modalities: From Vodcasts to Audio-Supported Reading
- Understanding gambling related harm: a proposed definition, conceptual framework, and taxonomy of harms
- Effects of sources of social support and resilience on the mental health of different age groups during the COVID-19 pandemic