Open-ended interview questions and saturation

Susan C. Weller, Ben Vickers, H. Russell Bernard, Alyssa Blackburn, Stephen P. Borgatti, Clarence C. Gravlee, Jeffrey C. JohnsonView original
OverviewBalancedlynda voice
How many interviews does a qualitative researcher actually need? That sounds like it should have a clear answer by now. Qualitative research has been around for decades, and the standard advice has barely changed: keep interviewing until you reach saturation, until nothing new appears. But here's the problem. That advice tells you almost nothing about when to stop or how hard to probe. It doesn't clarify whether saturation at 12 interviews and saturation at 120 interviews are the same thing. And it certainly doesn't explain what happens when you ask respondents three questions instead of thirty. Susan Weller and six colleagues just ran one thousand one hundred and forty-seven real interviews across twenty-eight different studies to find out whether that advice actually holds up, and the answer changes how you should design qualitative work from the start. The standard framing of saturation actually contains two distinct ideas that often get conflated. Theoretical saturation, the original concept from Glaser and Strauss's grounded theory, is a conceptual ideal: you stop when you've mapped the main ideas and their relationships well enough to build a theory. Thematic saturation is what you can observe directly in data: the accumulation of new items slows to a trickle. Most researchers treat these concepts interchangeably. Weller and colleagues argue they are different targets requiring different expectations, and that chasing the wrong one has real consequences. Prior empirical work hasn't settled the question. Guest and colleagues found twelve to sixteen interviews adequate in some settings. Statistical work by Fugard and Potts showed that for a theme present in twenty percent of the population, you need fourteen independent interviews for a ninety-five percent chance of seeing it at least once. Tran and colleagues estimated that samples beyond fifty would add less than one new theme per additional person. But those estimates depend heavily on how common the ideas are and on how much each respondent is allowed to say. That last variable is the one the field has mostly ignored. To get traction on all of this, Weller and colleagues assembled a large empirical base: twenty-eight free-listing datasets drawn from eighteen topical domains, totaling one thousand one hundred and forty-seven interviews. Free-listing means asking respondents to name everything they can think of within a topic — kinds of fruits, things mothers do, kinds of recreational drugs, racial and ethnic groups, birds, flowers. The domains ranged from tightly bounded to sprawling and open-ended, and datasets included between twenty and ninety-nine interviews each. The team fit generalized linear models — ordinary least squares, Poisson, and negative binomial — to describe how new items accumulate as sample size grows, and they applied capture-recapture methods to estimate total domain size. The capture-recapture logic is intuitive: if respondent one names fifteen items and respondent two names thirty-one, and they overlap on five, you estimate total domain size by multiplying fifteen by thirty-one and dividing by five. That gives you ninety-three, the product of their list lengths divided by the matches. The results on saturation itself are both concrete and sobering. When Weller and colleagues defined saturation as the point where fewer than one new item per additional person is expected, the median sample size across their twenty-eight examples was seventy-five, with a range of fifteen to one hundred and ninety-four. Using a slightly looser threshold of fewer than two new items per person, the median dropped to fifty, ranging from ten to one hundred and forty-six. Only five of the twenty-eight datasets actually reached the strict criterion within their original sample sizes. Two forces drive that enormous range. Domain size is the first. Fruits saturated at fifteen interviews; unbounded domains like things mothers do required an estimated one hundred and sixty-one. Saturation tended to occur at roughly ninety percent of the estimated total domain, so large domains push saturation far out almost by definition. The second force is more surprising. Saturation was reached more quickly when respondents gave fewer responses per person. When people provide only a few answers, they mostly supply the most common items, the accumulation curve drops sharply, and the threshold is crossed early. When respondents are probed deeply and give many responses, they contribute more salient items sooner, but they also keep adding rare, idiosyncratic items that create a long tail. That long tail is what pushes saturation out. So deep probing makes saturation harder to reach, even as it makes the resulting data more useful. That's the counterintuitive core of the finding. And that's exactly where the paper pivots. Weller and colleagues argue that chasing complete saturation is often the wrong goal. What researchers usually care about are the most salient items — the ideas that are widely shared, culturally central, and likely to show up again and again across respondents. Salience is measured by combining how often an item is mentioned with how early in a list it appears. The Smith and Sutrop indices do this in slightly different ways, but in these data, the two correlated with each other at an average Spearman correlation of 0.95, and both correlated strongly with simple frequency. They're essentially measuring the same thing. The question then becomes: how many interviews do you need to capture those salient items? With exhaustive listing—probing respondents to name everything they can—ten interviews captured ninety-five percent of the most salient ideas. Across the twenty-eight examples, eighteen captured one hundred percent and eight captured between ninety-five and ninety-nine percent of salient items within the first ten interviews. That is a striking result. A sample of ten people, probed thoroughly, gets you nearly everything the domain's most important ideas have to offer. But the qualifier matters enormously. When respondents were limited to just three responses each— a common situation in loosely structured open-ended interviews—ten interviews captured only fifty-three percent of those same salient items. With twenty interviews at three responses each, that figure climbed to just sixty-eight percent. Interviews that elicited more responses per person retrieved roughly fifty percent more domain items overall. Semantic cueing — repeating a respondent's prior answers and asking for more — increased yield by approximately the same amount. The practical implication runs directly against the standard advice. The field has focused on interviewing more people as the path to adequate data. Weller and colleagues' evidence says the better lever is probing more deeply within each interview. The number of responses per person predicts salient item capture better than the raw number of interviews, at least for the ideas researchers most care about. This leads to what the paper calls saturation in salience — a redefined target for qualitative sample size decisions. Instead of asking when no new ideas emerge at all, ask when the most salient ideas have been reliably identified. That reframing has three practical levers. Domain size is the first: small, bounded domains saturate quickly and need fewer interviews; large, open-ended domains require more. Probing depth is the second: exhaustive listing with intensive follow-up recovers salient content far more efficiently than shallow questioning. And raw sample size is third, but it matters less than the other two. A sample of fourteen gives a ninety-five percent chance of observing any item with a population prevalence of twenty percent; that's useful to know, but it assumes you're probing deeply enough to surface those items in the first place. For researchers working in bounded domains — clinical themes within a specific population, categories within a defined system — small, well-probed samples may be entirely adequate, and the field's instinct toward larger samples may be adding cost without adding insight. For genuinely unbounded domains or for research where rare items matter, larger samples and continued probing remain necessary. But even there, the decision should be driven by the target: what prevalence of item do you need to reliably detect — not by the vague instruction to keep going until things stop changing. The deeper point is that saturation has been doing too much work as a concept. It bundles together questions about domain size, probing quality, and the researcher's actual goals, and it answers none of them precisely. Weller and colleagues have unpacked those components, measured them across nearly one thousand two hundred interviews, and shown that the most consequential variable is the one most often overlooked: how hard you press each person to tell you everything they know. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

How many interviews does a qualitative researcher actually need? That sounds like it should have a clear answer by now. Qualitative research has been around for decades, and the standard advice has barely changed: keep interviewing until you reach saturation, until nothing new appears. But here's the problem. That advice tells you almost nothing about when to stop or how hard to probe. It doesn't clarify whether saturation at 12 interviews and saturation at 120 interviews are the same thing. And it certainly doesn't explain what happens when you ask respondents three questions instead of thirty. Susan Weller and six colleagues just ran one thousand one hundred and forty-seven real interviews across twenty-eight different studies to find out whether that advice actually holds up, and the answer changes how you should design qualitative work from the start. The standard framing of saturation actually contains two distinct ideas that often get conflated. Theoretical saturation, the original concept from Glaser and Strauss's grounded theory, is a conceptual ideal: you stop when you've mapped the main ideas and their relationships well enough to build a theory. Thematic saturation is what you can observe directly in data: the accumulation of new items slows to a trickle. Most researchers treat these concepts interchangeably. Weller and colleagues argue they are different targets requiring different expectations, and that chasing the wrong one has real consequences.

Prior empirical work hasn't settled the question. Guest and colleagues found twelve to sixteen interviews adequate in some settings. Statistical work by Fugard and Potts showed that for a theme present in twenty percent of the population, you need fourteen independent interviews for a ninety-five percent chance of seeing it at least once. Tran and colleagues estimated that samples beyond fifty would add less than one new theme per additional person. But those estimates depend heavily on how common the ideas are and on how much each respondent is allowed to say. That last variable is the one the field has mostly ignored. To get traction on all of this, Weller and colleagues assembled a large empirical base: twenty-eight free-listing datasets drawn from eighteen topical domains, totaling one thousand one hundred and forty-seven interviews. Free-listing means asking respondents to name everything they can think of within a topic — kinds of fruits, things mothers do, kinds of recreational drugs, racial and ethnic groups, birds, flowers. The domains ranged from tightly bounded to sprawling and open-ended, and datasets included between twenty and ninety-nine interviews each.

The team fit generalized linear models — ordinary least squares, Poisson, and negative binomial — to describe how new items accumulate as sample size grows, and they applied capture-recapture methods to estimate total domain size. The capture-recapture logic is intuitive: if respondent one names fifteen items and respondent two names thirty-one, and they overlap on five, you estimate total domain size by multiplying fifteen by thirty-one and dividing by five. That gives you ninety-three, the product of their list lengths divided by the matches. The results on saturation itself are both concrete and sobering. When Weller and colleagues defined saturation as the point where fewer than one new item per additional person is expected, the median sample size across their twenty-eight examples was seventy-five, with a range of fifteen to one hundred and ninety-four. Using a slightly looser threshold of fewer than two new items per person, the median dropped to fifty, ranging from ten to one hundred and forty-six. Only five of the twenty-eight datasets actually reached the strict criterion within their original sample sizes. Two forces drive that enormous range. Domain size is the first. Fruits saturated at fifteen interviews; unbounded domains like things mothers do required an estimated one hundred and sixty-one.

Saturation tended to occur at roughly ninety percent of the estimated total domain, so large domains push saturation far out almost by definition. The second force is more surprising. Saturation was reached more quickly when respondents gave fewer responses per person. When people provide only a few answers, they mostly supply the most common items, the accumulation curve drops sharply, and the threshold is crossed early. When respondents are probed deeply and give many responses, they contribute more salient items sooner, but they also keep adding rare, idiosyncratic items that create a long tail. That long tail is what pushes saturation out. So deep probing makes saturation harder to reach, even as it makes the resulting data more useful. That's the counterintuitive core of the finding. And that's exactly where the paper pivots. Weller and colleagues argue that chasing complete saturation is often the wrong goal. What researchers usually care about are the most salient items — the ideas that are widely shared, culturally central, and likely to show up again and again across respondents. Salience is measured by combining how often an item is mentioned with how early in a list it appears. The Smith and Sutrop indices do this in slightly different ways, but in these data, the two correlated with each other at an average Spearman correlation of 0.95, and both correlated strongly with simple frequency. They're essentially measuring the same thing.

The question then becomes: how many interviews do you need to capture those salient items? With exhaustive listing—probing respondents to name everything they can—ten interviews captured ninety-five percent of the most salient ideas. Across the twenty-eight examples, eighteen captured one hundred percent and eight captured between ninety-five and ninety-nine percent of salient items within the first ten interviews. That is a striking result. A sample of ten people, probed thoroughly, gets you nearly everything the domain's most important ideas have to offer. But the qualifier matters enormously. When respondents were limited to just three responses each— a common situation in loosely structured open-ended interviews—ten interviews captured only fifty-three percent of those same salient items. With twenty interviews at three responses each, that figure climbed to just sixty-eight percent. Interviews that elicited more responses per person retrieved roughly fifty percent more domain items overall. Semantic cueing — repeating a respondent's prior answers and asking for more — increased yield by approximately the same amount. The practical implication runs directly against the standard advice. The field has focused on interviewing more people as the path to adequate data. Weller and colleagues' evidence says the better lever is probing more deeply within each interview.

The number of responses per person predicts salient item capture better than the raw number of interviews, at least for the ideas researchers most care about. This leads to what the paper calls saturation in salience — a redefined target for qualitative sample size decisions. Instead of asking when no new ideas emerge at all, ask when the most salient ideas have been reliably identified. That reframing has three practical levers. Domain size is the first: small, bounded domains saturate quickly and need fewer interviews; large, open-ended domains require more. Probing depth is the second: exhaustive listing with intensive follow-up recovers salient content far more efficiently than shallow questioning. And raw sample size is third, but it matters less than the other two. A sample of fourteen gives a ninety-five percent chance of observing any item with a population prevalence of twenty percent; that's useful to know, but it assumes you're probing deeply enough to surface those items in the first place.

For researchers working in bounded domains — clinical themes within a specific population, categories within a defined system — small, well-probed samples may be entirely adequate, and the field's instinct toward larger samples may be adding cost without adding insight. For genuinely unbounded domains or for research where rare items matter, larger samples and continued probing remain necessary. But even there, the decision should be driven by the target: what prevalence of item do you need to reliably detect — not by the vague instruction to keep going until things stop changing. The deeper point is that saturation has been doing too much work as a concept. It bundles together questions about domain size, probing quality, and the researcher's actual goals, and it answers none of them precisely. Weller and colleagues have unpacked those components, measured them across nearly one thousand two hundred interviews, and shown that the most consequential variable is the one most often overlooked: how hard you press each person to tell you everything they know. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Social Sciences