Timing in turn-taking and its implications for processing models of language

Stephen C. Levinson, Francisco TorreiraView original
OverviewBalancedalloy voice
Let me start with a timing riddle that sits at the center of everyday talk. In ordinary conversation, when one person finishes, the next comes in astonishingly fast—about two-tenths of a second later. That's the "200 millisecond puzzle." The classic turn-taking rules from Sacks, Schegloff, and Jefferson give us the choreography: turns end at recognizable junctures, called transition relevance places, and simple rules decide who goes next. But here's the catch. Planning even a single spoken word typically takes on the order of six-tenths of a second. So how do we consistently hit that 200 millisecond handoff? The only way it works is if listeners don't just listen. They predict where the turn is going and start building their reply early, before the current turn has actually ended. Levinson and Torreira dug into this by instrumenting a large telephone corpus known as Switchboard to get precise about what overlap and timing really look like in the wild. Earlier studies were suggestive but coarse. Heldner and Edlund talked about overlaps in roughly 40 percent of transitions. Ten Bosch and colleagues saw that partial overlaps rose from about 20 percent face-to-face to 27 percent on the phone, and overall overlap incidence climbed from 44 to 52 percent across those settings. Helpful, yes, but not enough texture to test theories. The Switchboard reanalysis set out to supply that texture. Here's the big picture. Across roughly 38 hours of dyadic calls, 77 percent of the signal is one person talking, 19.2 percent is silence, and just 3.8 percent is both people at once. If you set the silent stretches aside, over 95 percent of the speech is produced by a single speaker at a time. That's the broad rhythm we all feel: overwhelmingly, one voice carries the floor. And when the floor does switch, about 30 percent of those transfers happen with some overlap rather than a clean gap, which means overlap is common but not the dominant mode. What's striking is how short those overlaps are. The most frequent, so-called "between-overlaps," where Speaker B comes in as Speaker A is finishing and takes the floor, cluster around a tenth of a second at their peak, with a median around two-tenths and an average around a quarter of a second. Three-quarters of them are under about four-tenths. The other class—"within-overlaps," where two interpausal units share a little time but don't flip the floor—tend to start within about a third of a second of the overlapped bit and rarely run long. In other words, overlap lives in blips, not battles. Function tells the same story. When Levinson and Torreira hand-coded a sample of overlaps, nearly three-quarters were backchannels or agreements—yeah, right, mhm. About a third happened within half a second of a likely transition point in the overlapped turn. And in all but five percent of cases, at least one obvious feature was in play: a prompt backchannel, a nearby transition place, a bit of silence, or a cue like disfluency or an abandoned start. Overlap isn't random chaos. It clusters where the structure of the talk points to it. Now, you could write this off as a peculiarity of adult phone calls, except the timing pressure shows up early in life. Caregiver-infant protoconversations start as loose alternations. At around three months, the average switch is a leisurely one and a half seconds. As the first year unfolds and before full language arrives, those gaps tighten toward eight-tenths of a second; toddlers are still around a second, and only gradually do children inch toward adult-like rapidity. If you track their eyes while they watch people talk, three-year-olds start shifting gaze in the gap as if they're anticipating the next speaker, with saccade latencies around three-tenths of a second. And when you strip away words and leave only the melody—the prosody—even younger infants use that to guess where a turn might end. The scaffold for projection is laid down early, and prosody is one of the first rungs. By adulthood, the bottleneck is production. Indefrey and Levelt's timeline lays out the gears. From deciding what to say to handing a motor plan to the vocal tract takes on the order of 600 milliseconds for a single word. Conceptual preparation alone takes about 175 milliseconds; then come steps like retrieving the word, building its sound shape, and packaging its syllables before phonetic encoding. Put two nouns together and you push toward three-quarters of a second or a bit more. A simple three-word phrase often sits near 900 milliseconds. And breathing is not just along for the ride. In spontaneous dialogue, inhalation before speaking typically stretches past half a second, and ultrasound studies see the tongue starting to shape the upcoming sound more than a tenth of a second before you hear it. Sometimes there are detectable adjustments almost half a second before acoustic release. The pipeline has depth, and it takes time. If listeners had to wait for a turn to finish, then start that whole pipeline, no one would ever hit a 200 millisecond gap. The brain's workaround is to overlap comprehension with planning. You see this in electrical measures. Magyari and colleagues found that when people listened to turns whose endings were predictable, markers of response preparation rose well before the actual end—on the order of a second or more ahead. Gisladottir and colleagues showed that in adjacency pairs—question then answer, greeting then greeting—you can pick up early signatures of which speech act is coming within the first few hundred milliseconds of hearing it. And Riest and colleagues saw that when people could predict an answer, they prepared similarly to when they outright knew it. Prediction and pre-launch are not just nice; they're routine. Let's talk theories for how the baton passes. One family, which you can think of as opportunistic floor-sharing, follows Sacks and colleagues in seeing turn-taking as a jointly managed system where overlaps are allowed but usually trimmed by local norms. The Switchboard numbers line up: most of the stream is one voice, overlap at transfers is common but short, and those overlaps tend to appear in principled places—backchannels, near possible completions, after tiny silences, when someone aborts a start. You don't need a perfect choreography to get that pattern, just shared expectations and sensitivity to the unfolding structure. A second family, associated with James Duncan's work, puts weight on explicit end signals—the current speaker owns the floor until they cede it. There are end signals, and they matter: final pitch movements, lengthening at the end of phrases, syntactic closure. But the corpus pushes back on a pure signaling story. There is plenty of overlap, and much of it occurs without clear, single markers of yield. It's also hard to square that view with the regularity of short positive offsets across very different contexts. Heldner and Edlund sharpen the point by arguing you shouldn't expect a clean target of "no gap, no overlap" in real talk. The data agree. Timing is precise but not perfect. The third family is the processing account. It starts from the production latencies we've been talking about—hundreds of milliseconds to more than a second—and the plain observation that conversational gaps rarely stretch beyond a couple of tenths. From there, the logic is straightforward. Comprehenders begin predicting well before a turn is done, start building a response as soon as the gist or speech act is clear, then hold articulation until a go-signal—often prosodic or syntactic—flashes near the end of the other person's turn. That's how you get modal response offsets around 100 to 300 milliseconds while still honoring the physics of planning speech. Levinson and Torreira try to knit these strands into a single timeline. Early in a turn, the listener is already modeling possible completions. Morphosyntax gives long-range clues: if a turn opens with "if" or "whether," it's not done until you hear the second clause. As the last half-second approaches, predictions sharpen from structure to words, and multiple cues accumulate that the turn is coming in for a landing—phrase-final lengthening, a characteristic melody, a syntactic wrap-up. That bundle functions as the go-signal. Meanwhile, production has been quietly moving forward in the background: conceptual content picked, likely lemmas retrieved, sounds assembled, a motor plan queued. Just before launch, the vocal apparatus readies—on the order of two-tenths of a second—so that when the landing lights turn green, articulation can roll. It's a dance between a predictive listener and a buffered speaker. They tether that model to concrete corpus texture. Switchboard is chopped into interpausal units—segments bounded by silences of at least 180 milliseconds—a total of 50,510 units averaging about 1.68 seconds, with a median around 1.23 seconds. Floor transfers come in two flavors: with a gap, or with a between-overlap where there's no silence between the voices. Another class, within-overlaps, stack units without changing who holds the floor. The counts are large—more than fourteen thousand gaps, more than six thousand between-overlaps, a few thousand within-overlaps—and the distributions are strongly skewed toward brevity. You don't need the figures to see the picture. Most overlaps are shorter than two syllables. A couple of timing asymmetries help lock this down. Inter-speaker gaps—the handoffs—are shorter than intra-speaker gaps, with the latter exceeding the former by on the order of 150 milliseconds. The modal handoff hovers between 100 and 200 milliseconds, which is right where you'd expect if a go-signal crops up near the end and the response was warmed up in buffer. And when handoffs miss, they tend to miss small: a backchannel nudges in, a bit of silence creeps between, or a disfluent hitch creates space. That's the hum of a robust system, not a brittle one. You might reasonably ask: does any of this matter beyond timing trivia? It does, because this is the ecological niche of language. Conversation is where language is learned, where it evolved, and where we use it most. The Switchboard metrics—average turns on the scale of 1.6 seconds, handoffs at 100 to 200 milliseconds, overlaps that are typically a quarter of a second and overwhelmingly backchannels—are not side notes. They define the constraints any theory of comprehension and production has to live inside. And they square with the cognitive facts. Long production lags require prediction. Early development shows the scaffold is in place before full words. Neural measures, including electroencephalography, known as EEG, reveal preparation rising long before a turn ends. There are live debates, and they're healthy. Heldner and Edlund are right to warn us off any fantasy of laser-perfect "no gap, no overlap" timing. Duncan and others have documented genuine end signals, and we shouldn't ignore them. But taken together, the corpus, the developmental trajectory, the EEG, and articulatory traces all point to the same center: prediction drives the system, and turn-final cues trigger release. Where does that leave the open questions? Three stand out. One is neural granularity: how exactly do comprehension and production processes couple in time, down to the hundreds of milliseconds, across different kinds of turns? A second is cross-linguistic variation: prosodic systems differ, syntactic packaging differs—does the go-signal look the same in a tone language as in English, or does the system lean on other cues? The third is long-range projection: how far ahead can listeners realistically forecast in complex structures, and how do they hedge those bets when a turn veers? Those questions are exciting because we now have the scaffolding to answer them. We can anchor cognitive models to conversational facts. We can treat overlap not as a nuisance but as a diagnostic. And we can remember that the 200 millisecond puzzle is not a bug in human language; it's the feature that keeps talk alive—your brain quietly racing the clock so you can meet someone else's thought in time.

Let me start with a timing riddle that sits at the center of everyday talk. In ordinary conversation, when one person finishes, the next comes in astonishingly fast—about two-tenths of a second later. That's the "200 millisecond puzzle." The classic turn-taking rules from Sacks, Schegloff, and Jefferson give us the choreography: turns end at recognizable junctures, called transition relevance places, and simple rules decide who goes next.

But here's the catch. Planning even a single spoken word typically takes on the order of six-tenths of a second. So how do we consistently hit that 200 millisecond handoff?

The only way it works is if listeners don't just listen. They predict where the turn is going and start building their reply early, before the current turn has actually ended.

Levinson and Torreira dug into this by instrumenting a large telephone corpus known as Switchboard to get precise about what overlap and timing really look like in the wild. Earlier studies were suggestive but coarse. Heldner and Edlund talked about overlaps in roughly 40 percent of transitions.

Ten Bosch and colleagues saw that partial overlaps rose from about 20 percent face-to-face to 27 percent on the phone, and overall overlap incidence climbed from 44 to 52 percent across those settings. Helpful, yes, but not enough texture to test theories. The Switchboard reanalysis set out to supply that texture.

Here's the big picture. Across roughly 38 hours of dyadic calls, 77 percent of the signal is one person talking, 19.2 percent is silence, and just 3.8 percent is both people at once. If you set the silent stretches aside, over 95 percent of the speech is produced by a single speaker at a time.

That's the broad rhythm we all feel: overwhelmingly, one voice carries the floor. And when the floor does switch, about 30 percent of those transfers happen with some overlap rather than a clean gap, which means overlap is common but not the dominant mode.

What's striking is how short those overlaps are. The most frequent, so-called "between-overlaps," where Speaker B comes in as Speaker A is finishing and takes the floor, cluster around a tenth of a second at their peak, with a median around two-tenths and an average around a quarter of a second. Three-quarters of them are under about four-tenths.

The other class—"within-overlaps," where two interpausal units share a little time but don't flip the floor—tend to start within about a third of a second of the overlapped bit and rarely run long. In other words, overlap lives in blips, not battles.

Function tells the same story. When Levinson and Torreira hand-coded a sample of overlaps, nearly three-quarters were backchannels or agreements—yeah, right, mhm. About a third happened within half a second of a likely transition point in the overlapped turn.

And in all but five percent of cases, at least one obvious feature was in play: a prompt backchannel, a nearby transition place, a bit of silence, or a cue like disfluency or an abandoned start. Overlap isn't random chaos. It clusters where the structure of the talk points to it.

Now, you could write this off as a peculiarity of adult phone calls, except the timing pressure shows up early in life. Caregiver-infant protoconversations start as loose alternations. At around three months, the average switch is a leisurely one and a half seconds.

As the first year unfolds and before full language arrives, those gaps tighten toward eight-tenths of a second; toddlers are still around a second, and only gradually do children inch toward adult-like rapidity. If you track their eyes while they watch people talk, three-year-olds start shifting gaze in the gap as if they're anticipating the next speaker, with saccade latencies around three-tenths of a second. And when you strip away words and leave only the melody—the prosody—even younger infants use that to guess where a turn might end.

The scaffold for projection is laid down early, and prosody is one of the first rungs.

By adulthood, the bottleneck is production. Indefrey and Levelt's timeline lays out the gears. From deciding what to say to handing a motor plan to the vocal tract takes on the order of 600 milliseconds for a single word.

Conceptual preparation alone takes about 175 milliseconds; then come steps like retrieving the word, building its sound shape, and packaging its syllables before phonetic encoding. Put two nouns together and you push toward three-quarters of a second or a bit more. A simple three-word phrase often sits near 900 milliseconds.

And breathing is not just along for the ride. In spontaneous dialogue, inhalation before speaking typically stretches past half a second, and ultrasound studies see the tongue starting to shape the upcoming sound more than a tenth of a second before you hear it. Sometimes there are detectable adjustments almost half a second before acoustic release. The pipeline has depth, and it takes time.

If listeners had to wait for a turn to finish, then start that whole pipeline, no one would ever hit a 200 millisecond gap. The brain's workaround is to overlap comprehension with planning. You see this in electrical measures.

Magyari and colleagues found that when people listened to turns whose endings were predictable, markers of response preparation rose well before the actual end—on the order of a second or more ahead. Gisladottir and colleagues showed that in adjacency pairs—question then answer, greeting then greeting—you can pick up early signatures of which speech act is coming within the first few hundred milliseconds of hearing it. And Riest and colleagues saw that when people could predict an answer, they prepared similarly to when they outright knew it. Prediction and pre-launch are not just nice; they're routine.

Let's talk theories for how the baton passes. One family, which you can think of as opportunistic floor-sharing, follows Sacks and colleagues in seeing turn-taking as a jointly managed system where overlaps are allowed but usually trimmed by local norms. The Switchboard numbers line up: most of the stream is one voice, overlap at transfers is common but short, and those overlaps tend to appear in principled places—backchannels, near possible completions, after tiny silences, when someone aborts a start.

You don't need a perfect choreography to get that pattern, just shared expectations and sensitivity to the unfolding structure.

A second family, associated with James Duncan's work, puts weight on explicit end signals—the current speaker owns the floor until they cede it. There are end signals, and they matter: final pitch movements, lengthening at the end of phrases, syntactic closure. But the corpus pushes back on a pure signaling story.

There is plenty of overlap, and much of it occurs without clear, single markers of yield. It's also hard to square that view with the regularity of short positive offsets across very different contexts. Heldner and Edlund sharpen the point by arguing you shouldn't expect a clean target of "no gap, no overlap" in real talk. The data agree. Timing is precise but not perfect.

The third family is the processing account. It starts from the production latencies we've been talking about—hundreds of milliseconds to more than a second—and the plain observation that conversational gaps rarely stretch beyond a couple of tenths. From there, the logic is straightforward.

Comprehenders begin predicting well before a turn is done, start building a response as soon as the gist or speech act is clear, then hold articulation until a go-signal—often prosodic or syntactic—flashes near the end of the other person's turn. That's how you get modal response offsets around 100 to 300 milliseconds while still honoring the physics of planning speech.

Levinson and Torreira try to knit these strands into a single timeline. Early in a turn, the listener is already modeling possible completions. Morphosyntax gives long-range clues: if a turn opens with "if" or "whether," it's not done until you hear the second clause.

As the last half-second approaches, predictions sharpen from structure to words, and multiple cues accumulate that the turn is coming in for a landing—phrase-final lengthening, a characteristic melody, a syntactic wrap-up. That bundle functions as the go-signal. Meanwhile, production has been quietly moving forward in the background: conceptual content picked, likely lemmas retrieved, sounds assembled, a motor plan queued.

Just before launch, the vocal apparatus readies—on the order of two-tenths of a second—so that when the landing lights turn green, articulation can roll. It's a dance between a predictive listener and a buffered speaker.

They tether that model to concrete corpus texture. Switchboard is chopped into interpausal units—segments bounded by silences of at least 180 milliseconds—a total of 50,510 units averaging about 1.68 seconds, with a median around 1.23 seconds. Floor transfers come in two flavors: with a gap, or with a between-overlap where there's no silence between the voices.

Another class, within-overlaps, stack units without changing who holds the floor. The counts are large—more than fourteen thousand gaps, more than six thousand between-overlaps, a few thousand within-overlaps—and the distributions are strongly skewed toward brevity. You don't need the figures to see the picture. Most overlaps are shorter than two syllables.

A couple of timing asymmetries help lock this down. Inter-speaker gaps—the handoffs—are shorter than intra-speaker gaps, with the latter exceeding the former by on the order of 150 milliseconds. The modal handoff hovers between 100 and 200 milliseconds, which is right where you'd expect if a go-signal crops up near the end and the response was warmed up in buffer.

And when handoffs miss, they tend to miss small: a backchannel nudges in, a bit of silence creeps between, or a disfluent hitch creates space. That's the hum of a robust system, not a brittle one.

You might reasonably ask: does any of this matter beyond timing trivia? It does, because this is the ecological niche of language. Conversation is where language is learned, where it evolved, and where we use it most.

The Switchboard metrics—average turns on the scale of 1.6 seconds, handoffs at 100 to 200 milliseconds, overlaps that are typically a quarter of a second and overwhelmingly backchannels—are not side notes. They define the constraints any theory of comprehension and production has to live inside. And they square with the cognitive facts.

Long production lags require prediction. Early development shows the scaffold is in place before full words. Neural measures, including electroencephalography, known as EEG, reveal preparation rising long before a turn ends.

There are live debates, and they're healthy. Heldner and Edlund are right to warn us off any fantasy of laser-perfect "no gap, no overlap" timing. Duncan and others have documented genuine end signals, and we shouldn't ignore them.

But taken together, the corpus, the developmental trajectory, the EEG, and articulatory traces all point to the same center: prediction drives the system, and turn-final cues trigger release.

Where does that leave the open questions? Three stand out. One is neural granularity: how exactly do comprehension and production processes couple in time, down to the hundreds of milliseconds, across different kinds of turns?

A second is cross-linguistic variation: prosodic systems differ, syntactic packaging differs—does the go-signal look the same in a tone language as in English, or does the system lean on other cues? The third is long-range projection: how far ahead can listeners realistically forecast in complex structures, and how do they hedge those bets when a turn veers?

Those questions are exciting because we now have the scaffolding to answer them. We can anchor cognitive models to conversational facts. We can treat overlap not as a nuisance but as a diagnostic.

And we can remember that the 200 millisecond puzzle is not a bug in human language; it's the feature that keeps talk alive—your brain quietly racing the clock so you can meet someone else's thought in time.

More in Arts and Humanities