Unaddressed participants’ gaze in multi-person interactionoptimizing recipiency

Judith Holler, Kobin H. KendrickView original
OverviewBalancedjames voice
Three people are talking — let's call them A, B, and C. A asks B a question. B is about to answer. Meanwhile, C, who has no role in this exchange, is already looking at B before A has finished speaking. Nobody told C to do that. There was no signal, no gesture pointing the way. Yet C's eyes moved at precisely the right moment, to precisely the right person, almost every time. The question this lecture answers is: how? The answer turns out to live inside the structure of the question itself — not its content, but its grammatical shape. Getting there requires understanding just how extraordinary ordinary conversation actually is. Here is the core tension. Quantitative studies show that the gap between one speaker finishing and the next beginning is typically somewhere between zero and two hundred milliseconds. Yet psycholinguistic research tells us that preparing even a simple utterance takes at least six hundred milliseconds. Those two numbers cannot both be true unless the next speakers begin planning their response well before the current turn ends — tracking its structure in real time, projecting where it's going, and launching their reply before the floor is technically open. Sacks, Schegloff, and Jefferson laid out the theoretical framework for this in nineteen seventy-four: conversation runs on a system of transition-relevance places, moments where a turn becomes recognizably complete and a handoff becomes possible. The cognitive machinery underneath that system has been a research target ever since. Eye gaze is one of the best windows onto that machinery. The human eye's visible white sclera makes gaze direction unusually salient — we can read where someone is looking from across a room. Studies using eye-tracking have shown that observers watching dialogues tend to fixate the current speaker and then shift toward the next speaker in anticipation. Foulsham and colleagues found that observers fixated the next speaker an average of one hundred and fifty milliseconds before that person began to speak. Keitel and colleagues found that fifty-four percent of gaze shifts to the next speaker happened within a window that started five hundred milliseconds before the current turn ended. These are striking numbers. But there's a catch: almost all of that evidence came from people watching scripted, pre-recorded dialogues — staged conversations with careful enunciation and gaps between turns averaging around nine hundred milliseconds. That raised an obvious question. Were these results showing us how human conversation actually works, or how people watch a television show? Judith Holler and Kobin Kendrick set out to answer that question. Their methodological move was direct: instead of showing people videos of conversations, they put eye-tracking equipment on people who were actually having one. Each recording session seated three native English speakers in a triangle, equidistant from one another, in a sound-proofed room. Each participant wore a head-mounted microphone and a pair of SMI eye-tracking glasses sampling at thirty frames per second. Three high-definition cameras captured the scene from different angles. Everything was synchronized into a single multimodal record with a time resolution of about forty-one milliseconds per frame. The conversations were spontaneous — casual talk between people who mostly already knew each other. The analysis focused specifically on question-response sequences between two of the three speakers, which meant the third person was momentarily "unaddressed" — a bystander, but a watching, listening, engaged one. Holler and Kendrick wanted to know exactly when that third person shifted their gaze from the speaker asking the question to the speaker about to answer it. They identified one hundred and five such sequences across seven conversations, coded gaze frame by frame, and measured the timing of each shift against two reference points: the actual end of the question turn, and the turn's first possible completion — the earliest moment within the turn where it became recognizably complete. That distinction is the heart of the study. A turn can become possibly complete before it actually ends. The classic example from Sacks and colleagues: "What is your last name" is already a grammatically complete question before the speaker adds the address term "Loraine." That addition point is the first possible completion — where grammatical, prosodic, and pragmatic cues align to make a handoff relevant — and it need not coincide with the turn's actual end. In Holler and Kendrick's data, fifty-four percent of question turns had at least one possible completion before the turn ended. The headline result: sixty percent of gaze shifts from current to next speaker occurred before the actual turn ended. Seventy-three percent were planned before it — accounting for the roughly two hundred milliseconds it takes to initiate an eye movement. So the unaddressed participants were not waiting for silence. They were moving their eyes while the question was still being spoken. But the more precise finding sharpens that picture considerably. When Holler and Kendrick measured gaze-shift timing not against the turn end but against the first possible completion, the distribution clustered right around that point. After correcting for the two-hundred-millisecond planning lag, the peak of covert gaze-shift initiation landed just forty milliseconds before the first possible completion. In other words, the moment the turn became recognizably complete was essentially the moment the unaddressed participant's brain began redirecting their gaze. The obvious alternative explanation is that participants were just reacting to the next speaker starting to talk. To rule that out, Holler and Kendrick ran a subset analysis. For sequences where the responder began speaking more than two hundred milliseconds after the first possible completion — cases where the response onset could not plausibly have triggered the gaze shift — the timing distribution barely moved. The mode shifted from one hundred and sixty milliseconds to one hundred and five milliseconds. The clustering around the first possible completion held. The gaze shift was not a reaction to hearing someone start speaking. It was a response to recognizing that the current turn had become complete. There was one interesting exception. In fifteen sequences where the response began before the first possible completion, gaze shifted earlier, tracking the early response onset — and a significant correlation between gaze-shift timing and response onset emerged in that subset, with a Spearman correlation of 0.23 and a p-value below 0.05. So an unusually early response can pull gaze forward ahead of schedule. But this is the minority case. The dominant driver is turn-structure recognition. This also explains something that had looked puzzling in earlier research. Studies using scripted dialogues, with their long average gaps and simple turn structures, often found that gaze shifts happened either in the gap between turns or at the very start of the next speaker's turn — not clearly anticipatory. Holler and Kendrick argue that when scripted stimuli collapse the distinction between first possible completion and actual turn end — because the turns are short and simple, with no extension after possible completion — you lose the ability to see what is actually driving gaze timing. Live, spontaneous conversation, with its richer and more varied turn structure, separates those two reference points, and the underlying pattern becomes visible. Holler and Kendrick call the organizing principle behind this timing "optimization of recipiency." It works like this. By shifting gaze at the first possible completion rather than earlier or later, the unaddressed participant accomplishes two things at once. They remain visually oriented to the current speaker for most of that speaker's turn — catching gestures, facial expressions, head movements, and all the bodily behavior that accompanies talk. And they arrive, eyes already directed at the next speaker, just as that person's response begins to unfold. They miss very little from either speaker. Recipiency, as Holler and Kendrick use the term, is not just about seeing — it is about displaying attention. Gaze is publicly readable. Looking at a speaker signals to everyone in the interaction that you are engaged, that you are a participant even when you are not speaking. The timing of gaze shifts manages that display across both the departing speaker and the incoming one. The unaddressed participant is not a passive bystander processing a stream of audio. They are actively calibrating their attention, frame by frame, to the micro-structure of the conversation around them. What this study ultimately shows is that the cognitive machinery of turn-taking runs deeper and more automatically than we tend to think. Without instruction, without effort, and without even being the person addressed, we track the grammatical trajectory of other people's sentences — and we move our eyes at the moment those sentences become complete. The coordination miracle of conversation turns out to rest, in part, on each participant doing this continuously, for every turn, whether they are speaking or not. C looks at B before A finishes asking the question because C already knows, from the shape of the sentence, that A is almost done. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Three people are talking — let's call them A, B, and C. A asks B a question. B is about to answer. Meanwhile, C, who has no role in this exchange, is already looking at B before A has finished speaking. Nobody told C to do that. There was no signal, no gesture pointing the way. Yet C's eyes moved at precisely the right moment, to precisely the right person, almost every time. The question this lecture answers is: how? The answer turns out to live inside the structure of the question itself — not its content, but its grammatical shape. Getting there requires understanding just how extraordinary ordinary conversation actually is. Here is the core tension. Quantitative studies show that the gap between one speaker finishing and the next beginning is typically somewhere between zero and two hundred milliseconds. Yet psycholinguistic research tells us that preparing even a simple utterance takes at least six hundred milliseconds.

Those two numbers cannot both be true unless the next speakers begin planning their response well before the current turn ends — tracking its structure in real time, projecting where it's going, and launching their reply before the floor is technically open. Sacks, Schegloff, and Jefferson laid out the theoretical framework for this in nineteen seventy-four: conversation runs on a system of transition-relevance places, moments where a turn becomes recognizably complete and a handoff becomes possible. The cognitive machinery underneath that system has been a research target ever since. Eye gaze is one of the best windows onto that machinery. The human eye's visible white sclera makes gaze direction unusually salient — we can read where someone is looking from across a room. Studies using eye-tracking have shown that observers watching dialogues tend to fixate the current speaker and then shift toward the next speaker in anticipation. Foulsham and colleagues found that observers fixated the next speaker an average of one hundred and fifty milliseconds before that person began to speak. Keitel and colleagues found that fifty-four percent of gaze shifts to the next speaker happened within a window that started five hundred milliseconds before the current turn ended. These are striking numbers.

But there's a catch: almost all of that evidence came from people watching scripted, pre-recorded dialogues — staged conversations with careful enunciation and gaps between turns averaging around nine hundred milliseconds. That raised an obvious question. Were these results showing us how human conversation actually works, or how people watch a television show? Judith Holler and Kobin Kendrick set out to answer that question. Their methodological move was direct: instead of showing people videos of conversations, they put eye-tracking equipment on people who were actually having one. Each recording session seated three native English speakers in a triangle, equidistant from one another, in a sound-proofed room. Each participant wore a head-mounted microphone and a pair of SMI eye-tracking glasses sampling at thirty frames per second. Three high-definition cameras captured the scene from different angles. Everything was synchronized into a single multimodal record with a time resolution of about forty-one milliseconds per frame. The conversations were spontaneous — casual talk between people who mostly already knew each other.

The analysis focused specifically on question-response sequences between two of the three speakers, which meant the third person was momentarily "unaddressed" — a bystander, but a watching, listening, engaged one. Holler and Kendrick wanted to know exactly when that third person shifted their gaze from the speaker asking the question to the speaker about to answer it. They identified one hundred and five such sequences across seven conversations, coded gaze frame by frame, and measured the timing of each shift against two reference points: the actual end of the question turn, and the turn's first possible completion — the earliest moment within the turn where it became recognizably complete. That distinction is the heart of the study. A turn can become possibly complete before it actually ends. The classic example from Sacks and colleagues: "What is your last name" is already a grammatically complete question before the speaker adds the address term "Loraine." That addition point is the first possible completion — where grammatical, prosodic, and pragmatic cues align to make a handoff relevant — and it need not coincide with the turn's actual end. In Holler and Kendrick's data, fifty-four percent of question turns had at least one possible completion before the turn ended.

The headline result: sixty percent of gaze shifts from current to next speaker occurred before the actual turn ended. Seventy-three percent were planned before it — accounting for the roughly two hundred milliseconds it takes to initiate an eye movement. So the unaddressed participants were not waiting for silence. They were moving their eyes while the question was still being spoken. But the more precise finding sharpens that picture considerably. When Holler and Kendrick measured gaze-shift timing not against the turn end but against the first possible completion, the distribution clustered right around that point. After correcting for the two-hundred-millisecond planning lag, the peak of covert gaze-shift initiation landed just forty milliseconds before the first possible completion. In other words, the moment the turn became recognizably complete was essentially the moment the unaddressed participant's brain began redirecting their gaze. The obvious alternative explanation is that participants were just reacting to the next speaker starting to talk. To rule that out, Holler and Kendrick ran a subset analysis. For sequences where the responder began speaking more than two hundred milliseconds after the first possible completion — cases where the response onset could not plausibly have triggered the gaze shift — the timing distribution barely moved.

The mode shifted from one hundred and sixty milliseconds to one hundred and five milliseconds. The clustering around the first possible completion held. The gaze shift was not a reaction to hearing someone start speaking. It was a response to recognizing that the current turn had become complete. There was one interesting exception. In fifteen sequences where the response began before the first possible completion, gaze shifted earlier, tracking the early response onset — and a significant correlation between gaze-shift timing and response onset emerged in that subset, with a Spearman correlation of 0.23 and a p-value below 0.05. So an unusually early response can pull gaze forward ahead of schedule. But this is the minority case. The dominant driver is turn-structure recognition. This also explains something that had looked puzzling in earlier research. Studies using scripted dialogues, with their long average gaps and simple turn structures, often found that gaze shifts happened either in the gap between turns or at the very start of the next speaker's turn — not clearly anticipatory. Holler and Kendrick argue that when scripted stimuli collapse the distinction between first possible completion and actual turn end — because the turns are short and simple, with no extension after possible completion — you lose the ability to see what is actually driving gaze timing.

Live, spontaneous conversation, with its richer and more varied turn structure, separates those two reference points, and the underlying pattern becomes visible. Holler and Kendrick call the organizing principle behind this timing "optimization of recipiency." It works like this. By shifting gaze at the first possible completion rather than earlier or later, the unaddressed participant accomplishes two things at once. They remain visually oriented to the current speaker for most of that speaker's turn — catching gestures, facial expressions, head movements, and all the bodily behavior that accompanies talk. And they arrive, eyes already directed at the next speaker, just as that person's response begins to unfold. They miss very little from either speaker. Recipiency, as Holler and Kendrick use the term, is not just about seeing — it is about displaying attention. Gaze is publicly readable. Looking at a speaker signals to everyone in the interaction that you are engaged, that you are a participant even when you are not speaking. The timing of gaze shifts manages that display across both the departing speaker and the incoming one. The unaddressed participant is not a passive bystander processing a stream of audio. They are actively calibrating their attention, frame by frame, to the micro-structure of the conversation around them.

What this study ultimately shows is that the cognitive machinery of turn-taking runs deeper and more automatically than we tend to think. Without instruction, without effort, and without even being the person addressed, we track the grammatical trajectory of other people's sentences — and we move our eyes at the moment those sentences become complete. The coordination miracle of conversation turns out to rest, in part, on each participant doing this continuously, for every turn, whether they are speaking or not. C looks at B before A finishes asking the question because C already knows, from the shape of the sentence, that A is almost done. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Arts and Humanities