Quantile regression estimands and modelsrevisiting the motherhood wage penalty debate

Nicolai T. Borgen, Andreas Haupt, Øyvind N. WiborgView original
HighlightsBalancedjames voice
Imagine a researcher running two quantile regression models on the same dataset — the same women, same wages, and the same motherhood variable. One model shows the penalty growing as you move up the wage distribution, while the other shows it shrinking. Both analyses are technically correct; they are just answering completely different questions. That's the problem that Borgen, Haupt, and Wiborg put on the table. The core issue is a mismatch between what researchers want to know and the model they select. There are two families of quantile regression, and they measure fundamentally different things. Quantile Treatment Effects, or QTEs, are individual-level estimands. Working in the potential outcomes framework — the language of what would have happened under a different treatment — a QTE asks how a specific person's wage rank would change if they became a mother, compared to if they didn't. Interpreting this requires an assumption called rank invariance, or the weaker rank similarity, meaning that individuals' expected ranks remain comparable across those two potential states. Unconditional quantile regression, or UQR, developed by Firpo and colleagues in 2009, asks something else entirely: how does a covariate shift the quantiles of the overall wage distribution across the entire population? UQR accomplishes this by constructing a recentered influence function, or RIF — which is a measure of how much each individual observation contributes to a given quantile of the full distribution. Because UQR compares a world where everyone is treated to one where no one is, its coefficients are sensitive to the share of treated individuals and to how dense the outcome distribution is at that quantile. The estimation toolkits differ accordingly. For QTEs, the standard approach is Firpo's 2007 two-step propensity-score method: model the probability of treatment, then run a conditional quantile regression weighted by the inverse of those probabilities. Frölich and Melly extended this to handle instrumental variables. Newer estimators, like Powell's generalized quantile regression and Borgen and colleagues' residualized quantile regression, handle continuous treatments and high-dimensional fixed effects. For UQR, the two-step RIF ordinary least squares procedure is the main method: compute the RIF for each observation, then regress it on covariates using ordinary least squares. The practical warning is clear — UQR was not built to estimate individual-level QTEs. Using it for that purpose produces incorrect answers. The simulations make this point evident. With a uniform individual treatment effect of four units and one-third of the sample treated, the overall mean rises by 1.33 — but the fifth percentile of the full distribution shifts by only 0.8, while the ninety-fifth percentile shifts by 2.0. More strikingly, when the treated share is just 10 percent, UQR coefficients decrease across the distribution; when it's 90 percent, they increase. These are opposite patterns, but with identical individual effects. In the real motherhood wage data, Borgen and colleagues found that UQR and QTE estimates actually aligned — a reassuring result for that literature. However, the simulations show this agreement is not guaranteed, and the conceptual stakes are clear: pick the wrong model, and you build the wrong theory. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Imagine a researcher running two quantile regression models on the same dataset — the same women, same wages, and the same motherhood variable. One model shows the penalty growing as you move up the wage distribution, while the other shows it shrinking. Both analyses are technically correct; they are just answering completely different questions. That's the problem that Borgen, Haupt, and Wiborg put on the table. The core issue is a mismatch between what researchers want to know and the model they select. There are two families of quantile regression, and they measure fundamentally different things. Quantile Treatment Effects, or QTEs, are individual-level estimands. Working in the potential outcomes framework — the language of what would have happened under a different treatment — a QTE asks how a specific person's wage rank would change if they became a mother, compared to if they didn't. Interpreting this requires an assumption called rank invariance, or the weaker rank similarity, meaning that individuals' expected ranks remain comparable across those two potential states. Unconditional quantile regression, or UQR, developed by Firpo and colleagues in 2009, asks something else entirely: how does a covariate shift the quantiles of the overall wage distribution across the entire population?

UQR accomplishes this by constructing a recentered influence function, or RIF — which is a measure of how much each individual observation contributes to a given quantile of the full distribution. Because UQR compares a world where everyone is treated to one where no one is, its coefficients are sensitive to the share of treated individuals and to how dense the outcome distribution is at that quantile. The estimation toolkits differ accordingly. For QTEs, the standard approach is Firpo's 2007 two-step propensity-score method: model the probability of treatment, then run a conditional quantile regression weighted by the inverse of those probabilities. Frölich and Melly extended this to handle instrumental variables. Newer estimators, like Powell's generalized quantile regression and Borgen and colleagues' residualized quantile regression, handle continuous treatments and high-dimensional fixed effects. For UQR, the two-step RIF ordinary least squares procedure is the main method: compute the RIF for each observation, then regress it on covariates using ordinary least squares. The practical warning is clear — UQR was not built to estimate individual-level QTEs. Using it for that purpose produces incorrect answers.

The simulations make this point evident. With a uniform individual treatment effect of four units and one-third of the sample treated, the overall mean rises by 1.33 — but the fifth percentile of the full distribution shifts by only 0.8, while the ninety-fifth percentile shifts by 2.0. More strikingly, when the treated share is just 10 percent, UQR coefficients decrease across the distribution; when it's 90 percent, they increase. These are opposite patterns, but with identical individual effects. In the real motherhood wage data, Borgen and colleagues found that UQR and QTE estimates actually aligned — a reassuring result for that literature. However, the simulations show this agreement is not guaranteed, and the conceptual stakes are clear: pick the wrong model, and you build the wrong theory. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Mathematics