Information structure in discourseTowards an integrated formal theory of pragmatics
Imagine every conversation you care about as a shared investigation. Not just chatter, not a pile of sentences, but a coordinated search for how things are. That's the spirit Craige Roberts channels, building on Grice, Lewis, and the planning insights of Grosz and Sidner.
The guiding picture is simple and powerful: discourse is a game of inquiry. There are goals we're trying to reach, rules we honor to get there, moves we make along the way, and strategies that tell us which move to try next. Some moves are brilliant. Some are just okay. And yes, luck matters.
To keep score in that game, you need a scoreboard. Roberts calls it the information structure of the discourse. Think of Stalnaker's common ground—the shared assumptions that define our live possibilities—and imagine pushing that context toward a single winner: the actual world.
That's the Big Question animating everything. Conventional rules like syntax and compositional semantics set the bounds. Conversational rules—Grice's maxims of relevance and cooperation—steer the play.
The organizing principle is crisp: a move is relevant if it helps answer the question we're currently discussing, either by offering a partial answer or by teeing up a subquestion that would get us closer.
How do we model that state, concretely? Roberts lays out a tuple—call it the InfoStr state—that bundles the moving parts. There's a set of moves, both setup moves like questions and payoff moves like assertions.
There's a set of questions, each denoting a collection of possible answers. There's a set of assertions, each ruling out worlds where the assertion would be false. There's an ordering of moves because timing matters.
There’s the set of accepted moves, the ones the participants have actually taken on board. There’s a function that tells us, for each move, what the common ground was just before it. And there’s the stack of questions under discussion—the question under discussion—that acts like a push-down store.
When we accept a question, it goes on top. When it's answered, or we decide we can't answer it right now, we pop it off.
This stack gives us discipline. The top question is the one you owe relevance to. If you assert something, it should at least partly answer that question.
If you ask something, it should be part of a plan to answer it. The plan is explicit. Roberts calls it a strategy of inquiry for a question q, and she treats it as an ordered pair: a strategy label plus a set of subquestions.
In plain talk, a strategy doesn't just say "let's find out who did what." It says "first, we'll figure out who the players are, then what each person did, and the answers to those smaller questions will together settle the big one."
You can hear the logic in a toy dialogue. Suppose the top question is, "Who ate what?" You split it into two subproblems—call them a and b—and each splits again: two little "who" bits, two little "what" bits. Answering the two subquestions under a gives you a partial answer to a.
Doing the same for b gives you the rest. Once a and b are both resolved, the common ground now settles the top question, and you pop it from the stack. The magic is in the entailment: complete answers to lower subquestions are built so they entail, and thereby support, partial answers to higher ones. It's a scaffolded climb.
Under the hood, questions have a clean semantics that makes this work. Following Hamblin, and later Groenendijk and Stokhof, a question denotes a set of alternatives—the question alternative set. That set partitions the live possibilities into cells, one for each complete answer.
A partial answer is a proposition that lets you evaluate at least one element in that set. A complete answer settles every element. Accept a complete answer, and you throw out the other cells. The scoreboard shrinks. Progress.
Make it concrete. Picture a tiny world with Mary, Alice, and Grace, and the relation invite. Ask, "Who did Mary invite?" The alternatives carve the space into four cells: Mary invited both Alice and Grace;
Mary invited Alice but not Grace; Mary invited Grace but not Alice; and Mary invited nobody. Saying "Mary didn't invite Grace" kicks out two cells but leaves open whether she invited Alice.
It's a partial answer—legitimately relevant under the top question, and precisely so.
Now for the move that gives this theory its bite: prosody isn't decoration. It's a commitment. Roberts argues that English intonational focus is presuppositional—your prosody says, "I'm answering this kind of question," and it demands that the context cooperate.
She implements that with focus marking—call it F-marking—on the focused bit of an utterance. Then she defines the focus alternative set for that marked phrase in a way that listeners can track: replace each focused or wh word with a variable, and consider the meanings you get by plugging in different values for those variables. That set is what your intonation points to.
Here's the crucial alignment condition. A move is congruent to a particular question if and only if the move's focus alternative set matches the question's alternatives. In other words, your intonation advertises a space of contrast; the current question provides exactly that space; when those two line up, the move is felicitous.
Misalign them, and you feel it as infelicity or oddness. Roberts elevates that to a presupposition: a prosodically focused assertion presupposes that it's a congruent answer to the question under discussion. Extend that across moods, and you get a general rule—any focus-bearing utterance presupposes congruence with the active question.
This is not just a slogan. It gives a sharp test: when the focus pattern doesn't fit the question, the move solicits accommodation—it pressures the discourse to adopt a matching question. Sometimes that works, especially when the context is pliable.
Sometimes it doesn't, and the result is a corrective or a "wait, what are we talking about?" moment. Either way, the information structure spine—common ground and question under discussion stack—does the bookkeeping.
What about the sound system that carries those commitments? Roberts leans on the modern intonational toolkit that grew out of work by Jackendoff and others. Each utterance has at least one intonational phrase.
Inside, there's at least one focused constituent and at least one pitch accent, and every accent lands inside a focused bit. There's a phrase accent and a boundary tone per phrase. The last pitch accent in the focused constituent bears nuclear stress—our old nuclear stress rule.
Two classic contours, often called A and B, leverage different boundary tones to signal different focus structures. B contours use a boundary sequence like rising then high—think low high percent—to mark an independent focus, essentially introducing an argument slot of the presupposed question. A contours, often with a high boundary, cue a dependent focus, one that fills that slot.
As you stack subquestions, intermediate phrases help weight multiple focused elements—more than a single nuclear stress could carry alone. The point isn't the label soup. It's that prosody systematically encodes the very contrasts the question under discussion organizes, and the presuppositional link forces them to match.
One place this pays off is with focus-sensitive items like only. There's a long tradition, especially in Rooth's alternative semantics, of deriving only's domain from the focus it takes scope over, computing a set of properties and quantifying over them. Roberts argues you get better predictions if you put the domain restriction where the action already is: in the pragmatics of the question under discussion and the presuppositions carried by prosodic focus.
Take, "Mary only invited Lyn for dinner." What's the hidden question this sentence is presupposing and aiming to answer? Not "what properties does Mary have," full stop—that's too broad. The presupposed inquiry is much tighter: which individual or individuals are such that Mary's inviting them for dinner exhausts the relevant possibilities?
For the sentence to be a sensible move, the domain of only must be limited to the set of inviting relations at issue under the current question—typically, "Who did Mary invite for dinner?" If you instead let only range over a big Rooth-style set of unrelated properties, you predict weird truth conditions and miss the felt felicity judgments. Kadmon's classic examples show how those broad domains can misfire. Under Roberts's account, the question under discussion and the focus pattern work together to serve the live strategy of inquiry, and the domain of only falls out of that relevance constraint—no extra lexical machinery needed.
Prosody again does teamwork here. Those A and B contours, with their distinct boundary tones, help mark whether a focus is introducing a new argument slot in the presupposed question or filling one that's already on the stack. That matters when you're stepping through a multi-part plan—independent focus beckons a subquestion; dependent focus supplies a partial answer.
It shows up in more playful corners of language, too. In a question like "Do you want coffee or tea?" it can be used metalinguistically to present the union of independently focused alternatives, with prosody presupposing a superquestion that bundles them together. You can get a conjoined feel without invoking a separate mechanism of conjunction reduction.
The contours do the signaling; the question under discussion stack does the structuring.
Let's pull the pieces together. The integrated theory gives you a single, disciplined scoreboard. When a move is accepted, the common ground updates: assertions intersect the shared possibilities, throwing out worlds where the assertion would be false; questions, by contrast, don't prune worlds; they push onto the question under discussion stack and commit the participants to answering them.
A move counts as relevant just when it either contributes a partial answer to the top question or advances a recognized strategy by introducing a licensed subquestion. As complete answers arrive, the partition induced by the question collapses to the winning cell, the question pops off, and attention returns to whatever remains on the stack. Throughout, prosody presupposes that the focus it carries is congruent with that very stack, and the match—or mismatch—guides interpretation, accommodation, or repair in real time.
Two strengths stand out. First, questions as alternatives semantics does the heavy lifting for both answerhood and focus. You don't need to posit dedicated lexical focus operators to explain why focus-sensitive expressions behave; you get them for free from the interplay of alternatives, relevance, and presupposition.
Second, the strategy of inquiry layer—the explicit plan of subquestions—explains why discourse feels goal-directed rather than episodic. It's not just a pile of moves; it's a staircase whose steps entail support for the landings above them.
There are open fronts, of course. Presuppositions project in messy ways across complex sentences and across turns, and a full dynamic update theory has to account for how interlocutors repair, cancel, or clarify without collapsing into contradiction. Roberts treats cancellation less as a primitive and more as a move that adjusts the plan post hoc—think of it as an explicit reshuffling of the question under discussion stack and the commitments tied to it.
That's a design choice that keeps the architecture lean, but it invites careful empirical work.
If you're a linguist, there's an obvious invitation to go cross-linguistic. How do tone languages or languages with morphological focus marking align prosodic or segmental cues with strategy of inquiry constraints? If you're a builder of dialogue systems, there's another.
The InfoStr tuple, the question under discussion push-down stack, and the congruence presupposition are exactly the scaffolds you want in a system that aims to be helpful: know the current question, commit to a plan of subquestions, and make every move relevant.
But even if you're just listening on your way to work, here's the gist to carry with you. Conversations feel good when they move. They feel coherent when each turn either chips away at the question we're actually asking, or sets up the next chip with a smaller, sharper question.
Prosody isn't window dressing; it's a promise about the plan. And the scoreboard—the shared, evolving common ground and the stack of questions we've pledged to pursue—lets us keep that promise, together.
Related lectures
- From mission to market: a case study and analysis of the commercialisation of institutional publishing
- Exploring Intersectionality: Black Female Identities and Cultural Performance
- Lexical Alternatives as a Source of Pragmatic Presuppositions
- Positive and negative emotions underlie motivation for L2 learning
- Practicing a Musical Instrument in Childhood is Associated with Enhanced Verbal Ability and Nonverbal Reasoning
- Cultural Locations of Disability