The impact of non-response bias due to sampling in public health studiesA comparison of voluntary versus mandatory recruitment in a Dutch national survey on adolescent health

Kei Long Cheung, Peter M. ten Klooster, Cees Smit, Hein de Vries, Marcel E. PieterseView original
OverviewBalancedlynda voice
Every public health number you've ever heard about teen drinking, smoking, or sexual behavior came from a survey. Surveys only count the people who show up. That sounds obvious, but buried in that obvious fact is a problem that can quietly warp everything. The teenagers least likely to mail back a questionnaire are precisely the teenagers most likely to be doing the things the questionnaire tries to measure. So, the headline statistic and the true prevalence might be drifting apart — systematically and invisibly — every time a low-response survey goes out. That's the problem Cheung and colleagues set out to quantify. Their core theoretical claim is simple: voluntary participation preferentially selects lower-risk respondents, which drives prevalence estimates downward. They had a rare opportunity to test it cleanly. In two thousand eleven, the same adolescent health questionnaire — the Electronic Monitor and Health Education, abbreviated as E-MOVO — was administered to young people in two geographically adjacent, demographically similar regions of the Netherlands: Twente and IJsselland. The instrument covered tobacco use, alcohol, mental health, sexual behavior, subjective health, and school experience. Same questions. Same year. Same country. But the recruitment methods were completely different. In Twente, participation was essentially mandatory: nine thousand three hundred sixty students completed the survey during a single class session. In IJsselland, a voluntary postal mailing invited one thousand nine hundred fifty-two young people to complete the same questionnaire online using a personalized code. This was a natural experiment. The researchers could compare two populations that should look similar and ask: how much does the recruitment method alone change what you find? The first thing it changes is who shows up. The voluntary IJsselland sample was fifty-five point eight percent female, compared to forty-nine point two percent in the mandatory Twente sample. The voluntary sample was also skewed toward higher education — sixty-one point three percent in the higher academic track versus fifty-one point five percent in the mandatory group. These gaps matter because females, older respondents, and more educated respondents are already known to return postal questionnaires at higher rates, and they tend to report healthier behaviors. So, the voluntary sample was stacked demographically toward lower-risk profiles before a single health question was even answered. Here's where the study gets interesting. Cheung and colleagues didn't just show that the two samples looked different. They ran logistic regression models that adjusted for gender, age, and education — essentially asking: once you control for who showed up, does the sampling method itself still drive a health gap? It did. The prevalence differences between the two samples were not small. Daily smoking was nine point one percent in the mandatory Twente group and three point five percent in the voluntary IJsselland group. Ever having consumed alcohol was fifty-one point eight percent versus twenty-six point two percent. Alcohol use in the past four weeks was forty-one point five percent versus eleven point one percent. That last figure is roughly a four-fold difference — and it survived demographic adjustment. The multivariable odds ratio for recent alcohol use, comparing voluntary to mandatory, was zero point one three, with a confidence interval of zero point one one to zero point one six. That means voluntary respondents had about one-eighth the odds of reporting recent drinking, even after accounting for who they were demographically. For smoking, the adjusted odds ratio was zero point four five. For ever having had sexual intercourse, it was zero point seven six. For mental health, classified using the twenty-five item Strengths and Difficulties Questionnaire, eight point one percent of the mandatory sample fell in the "possible" concern category versus five point five percent in the voluntary group — and that gap also persisted after adjustment. To put this in context, Cheung and colleagues compared these adolescent gaps to regional adult survey data from the same period. Adult smoking rates between the two regions were twenty-three point nine percent versus twenty-two point zero percent — a relative risk of about one point zero eight. Among the adolescents, the same comparison produced a relative risk of two point six. The adolescent gap far exceeds anything that could be explained by genuine regional differences in behavior. The gap is a measurement artifact. It's what happens when your recruitment method selects for the people least likely to be doing what you're studying. The implication is uncomfortable for public health monitoring. If agencies rely on voluntary surveys to track trends in adolescent substance use or sexual behavior, they are likely underestimating both the prevalence and, potentially, changes over time. The problem is topic-specific: people may decline a survey precisely because the topic is sensitive. Standard demographic weighting won't fully correct for that because the bias isn't just about age or gender — it's about motivation and behavior. Now, here's the finding that adds real practical nuance. While the prevalence estimates were badly skewed by sampling method, the relationships between variables were not. Cheung and colleagues tested this directly, constructing interaction terms between sampling method and each health variable and following the Baron and Kenny procedure for moderation analysis. They found no support for any moderating effect. Tobacco use related to alcohol use in the same way across both samples. Mental health related to school experience in the same way. The internal architecture of how these variables connect to each other — the structure of associations — was stable regardless of how people were recruited. This is a meaningful distinction. If you're using survey data to study mechanisms — how risk factors cluster together, which behaviors predict which outcomes — voluntary surveys may serve you reasonably well. The correlations hold. But if you're trying to say how many teenagers in your region are drinking every week or what percentage smoked last month, a voluntary survey could give you a number that's dramatically too low. Researchers studying within-subject associations can have some confidence in their findings from voluntary samples, while researchers making prevalence claims should be much more cautious. The study has limitations worth naming. It is cross-sectional, so it can't track change over time. It was conducted in two specific Dutch regions, and the assumption that Twente and IJsselland are truly comparable — that the mandatory sample had minimal non-response at the individual level — can't be fully verified. Cheung and colleagues acknowledge that not all sources of bias can be entirely disentangled. The mandatory sample's near-complete participation was achieved at the cluster level — by schools agreeing to participate — so some selection still happened, just earlier in the process. Still, the core lesson holds. The paper recommends that public health researchers maximize response rates in voluntary studies and apply analytic techniques to correct for non-response bias when estimating prevalence. Agencies monitoring adolescent health trends should be aware that low-response surveys systematically remove the higher-risk individuals whose absence reshapes the numbers. The way you recruit changes what you find. That's not a methodological footnote. It's the whole story. The teenagers who don't mail back the survey — who lose it, ignore it, or quietly decide they'd rather not answer questions about last weekend — are the teenagers whose behavior your public health system most needs to understand. The voluntary survey, as a design, makes them disappear. It replaces them with a sample that's more female, more educated, more likely to have had a good week at school, and considerably less likely to have had a drink. The numbers that come out the other end look cleaner than reality. Policy built on cleaner-than-reality numbers will always underestimate the scale of what it's trying to fix. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Every public health number you've ever heard about teen drinking, smoking, or sexual behavior came from a survey. Surveys only count the people who show up. That sounds obvious, but buried in that obvious fact is a problem that can quietly warp everything. The teenagers least likely to mail back a questionnaire are precisely the teenagers most likely to be doing the things the questionnaire tries to measure. So, the headline statistic and the true prevalence might be drifting apart — systematically and invisibly — every time a low-response survey goes out. That's the problem Cheung and colleagues set out to quantify. Their core theoretical claim is simple: voluntary participation preferentially selects lower-risk respondents, which drives prevalence estimates downward. They had a rare opportunity to test it cleanly. In two thousand eleven, the same adolescent health questionnaire — the Electronic Monitor and Health Education, abbreviated as E-MOVO — was administered to young people in two geographically adjacent, demographically similar regions of the Netherlands: Twente and IJsselland. The instrument covered tobacco use, alcohol, mental health, sexual behavior, subjective health, and school experience. Same questions. Same year. Same country. But the recruitment methods were completely different.

In Twente, participation was essentially mandatory: nine thousand three hundred sixty students completed the survey during a single class session. In IJsselland, a voluntary postal mailing invited one thousand nine hundred fifty-two young people to complete the same questionnaire online using a personalized code. This was a natural experiment. The researchers could compare two populations that should look similar and ask: how much does the recruitment method alone change what you find? The first thing it changes is who shows up. The voluntary IJsselland sample was fifty-five point eight percent female, compared to forty-nine point two percent in the mandatory Twente sample. The voluntary sample was also skewed toward higher education — sixty-one point three percent in the higher academic track versus fifty-one point five percent in the mandatory group. These gaps matter because females, older respondents, and more educated respondents are already known to return postal questionnaires at higher rates, and they tend to report healthier behaviors. So, the voluntary sample was stacked demographically toward lower-risk profiles before a single health question was even answered.

Here's where the study gets interesting. Cheung and colleagues didn't just show that the two samples looked different. They ran logistic regression models that adjusted for gender, age, and education — essentially asking: once you control for who showed up, does the sampling method itself still drive a health gap? It did. The prevalence differences between the two samples were not small. Daily smoking was nine point one percent in the mandatory Twente group and three point five percent in the voluntary IJsselland group. Ever having consumed alcohol was fifty-one point eight percent versus twenty-six point two percent. Alcohol use in the past four weeks was forty-one point five percent versus eleven point one percent. That last figure is roughly a four-fold difference — and it survived demographic adjustment. The multivariable odds ratio for recent alcohol use, comparing voluntary to mandatory, was zero point one three, with a confidence interval of zero point one one to zero point one six. That means voluntary respondents had about one-eighth the odds of reporting recent drinking, even after accounting for who they were demographically. For smoking, the adjusted odds ratio was zero point four five. For ever having had sexual intercourse, it was zero point seven six.

For mental health, classified using the twenty-five item Strengths and Difficulties Questionnaire, eight point one percent of the mandatory sample fell in the "possible" concern category versus five point five percent in the voluntary group — and that gap also persisted after adjustment. To put this in context, Cheung and colleagues compared these adolescent gaps to regional adult survey data from the same period. Adult smoking rates between the two regions were twenty-three point nine percent versus twenty-two point zero percent — a relative risk of about one point zero eight. Among the adolescents, the same comparison produced a relative risk of two point six. The adolescent gap far exceeds anything that could be explained by genuine regional differences in behavior. The gap is a measurement artifact. It's what happens when your recruitment method selects for the people least likely to be doing what you're studying. The implication is uncomfortable for public health monitoring. If agencies rely on voluntary surveys to track trends in adolescent substance use or sexual behavior, they are likely underestimating both the prevalence and, potentially, changes over time. The problem is topic-specific: people may decline a survey precisely because the topic is sensitive. Standard demographic weighting won't fully correct for that because the bias isn't just about age or gender — it's about motivation and behavior.

Now, here's the finding that adds real practical nuance. While the prevalence estimates were badly skewed by sampling method, the relationships between variables were not. Cheung and colleagues tested this directly, constructing interaction terms between sampling method and each health variable and following the Baron and Kenny procedure for moderation analysis. They found no support for any moderating effect. Tobacco use related to alcohol use in the same way across both samples. Mental health related to school experience in the same way. The internal architecture of how these variables connect to each other — the structure of associations — was stable regardless of how people were recruited. This is a meaningful distinction. If you're using survey data to study mechanisms — how risk factors cluster together, which behaviors predict which outcomes — voluntary surveys may serve you reasonably well. The correlations hold. But if you're trying to say how many teenagers in your region are drinking every week or what percentage smoked last month, a voluntary survey could give you a number that's dramatically too low. Researchers studying within-subject associations can have some confidence in their findings from voluntary samples, while researchers making prevalence claims should be much more cautious.

The study has limitations worth naming. It is cross-sectional, so it can't track change over time. It was conducted in two specific Dutch regions, and the assumption that Twente and IJsselland are truly comparable — that the mandatory sample had minimal non-response at the individual level — can't be fully verified. Cheung and colleagues acknowledge that not all sources of bias can be entirely disentangled. The mandatory sample's near-complete participation was achieved at the cluster level — by schools agreeing to participate — so some selection still happened, just earlier in the process. Still, the core lesson holds. The paper recommends that public health researchers maximize response rates in voluntary studies and apply analytic techniques to correct for non-response bias when estimating prevalence. Agencies monitoring adolescent health trends should be aware that low-response surveys systematically remove the higher-risk individuals whose absence reshapes the numbers. The way you recruit changes what you find. That's not a methodological footnote. It's the whole story.

The teenagers who don't mail back the survey — who lose it, ignore it, or quietly decide they'd rather not answer questions about last weekend — are the teenagers whose behavior your public health system most needs to understand. The voluntary survey, as a design, makes them disappear. It replaces them with a sample that's more female, more educated, more likely to have had a good week at school, and considerably less likely to have had a drink. The numbers that come out the other end look cleaner than reality. Policy built on cleaner-than-reality numbers will always underestimate the scale of what it's trying to fix. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Social Sciences