Embedding epistemic modals in EnglishA corpus-based study

Valentine Hacquard, Alexis WellwoodView original
OverviewBalancedharper voice
The word "might" in the sentence "John thinks it might rain" either means something or it doesn't. That sounds like a trivially simple question, but it isn't. For decades, linguists have split into two camps over exactly this issue, and the answer has consequences for how we think about meaning itself. One side argues that epistemic modals — words like "might," "must," and "can" when they express what is possible or necessary given what we know — are genuine contributors to truth conditions. They are part of what makes a sentence true or false. The other side claims they float above the sentence like a tone of voice, marking the speaker's attitude rather than shaping the proposition. The debate has been hard to settle because the evidence supports both sides. Hacquard and Wellwood decided to stop arguing with constructed examples and to examine how the language actually behaves. The theoretical stakes are worth considering for a moment. In the Kratzerian tradition — named for the semanticist Angelika Kratzer — epistemic modals are treated as quantifiers over possible worlds. The sentence "John must be the murderer" is true if, in all worlds compatible with what is known, John is the murderer. That's a real semantic contribution. The rival view treats epistemic expressions as speaker assessments — evidential markers that report the speaker's degree of commitment without changing the underlying proposition. The diagnostic that both camps have leaned on is embedding. If epistemic modals can appear inside the scope of other operators — within questions, in the "if" clause of a conditional, or in the complement of a verb like "think" or "believe" — then they are behaving like ordinary semantic operators. If they cannot, that suggests they exist outside the sentence's meaning, at the speaker level. Classical accounts claimed that epistemics generally do not embed, but that claim turned out to be harder to defend than it appeared. Hacquard and Wellwood drew their data from the New York Times section of the English Gigaword Corpus, a massive written corpus that they parsed into fifteen million, six hundred ninety-one thousand, eight hundred fifty-nine sentences. Across that data, they found one hundred forty-nine thousand, two hundred nineteen tokens of "might," eighty-eight thousand, eight hundred fifty-nine of "must," and four hundred seventy-five thousand, five hundred ninety of "can." They examined three embedding environments: antecedents of conditionals, questions including embedded questions, and complements of attitude predicates. For "must" and "have to," they hand-annotated whether each token was epistemic or root, using paraphrases as guides — epistemic "must" reads as "it is probable that," while root "must" indicates an obligation. Interannotator agreement was high, with a kappa of 0.84. The parsed corpus allowed systematic comparison across environments at a scale that no constructed-example approach could match. Now to what they actually found — and this is where the picture becomes genuinely interesting. Starting with conditional antecedents, the "if" clause in "If it might rain, bring an umbrella" turns out to be nearly hostile to epistemic modals in real data. Only thirty tokens of "might" occurred in antecedents of conditionals — that's 0.02 percent of all "might" tokens in the corpus. "Can," by contrast, appeared there nine thousand, two hundred ninety-two times, or one point ninety-five percent of its corpus distribution. For "must," there were two hundred thirteen total tokens in "if" clauses, and in all but one case, the modal received a root interpretation. Epistemic "must" in conditional antecedents was effectively absent. That's not a categorical ban from theory — it's an empirical near-zero from actual language use. Questions tell a more complicated story. In matrix questions — questions sitting at the top of a sentence — "can" is far more common than "might": three point seventy-eight percent of "can" tokens versus zero point thirty-five percent of "might" tokens appear in matrix questions. That asymmetry fits the speaker-oriented picture; asking a question about your own uncertainty feels strange. However, embedded questions — questions embedded inside a larger sentence — do not show the same scarcity. "Might" appears in embedded questions at one point fifteen percent of its distribution, "can" at zero point eighty percent, and eighty-six cases of epistemic "must" are attested there. In complements of inquisitive verbs specifically, zero point ninety-two percent of "might" tokens appear, versus only zero point zero four percent of "can." So the ban is not on embedding in questions generally; it's more targeted than that. The richest evidence comes from attitude predicates — verbs like "think," "believe," "say," "want," and "order." This is where the data push back against a simple story. Epistemic modals do appear in the complements of some attitude verbs — abundantly, in fact, under verbs of acceptance, perception, certainty, and conjecture. However, under desideratives and directives — "want," "order," "demand" — they're essentially absent. The paper reports no instances of epistemic "have to" in desiderative or directive complements. Additionally, there's a revealing contrast within the attitude class: finite complements can host epistemic readings, while infinitival "have to" completely lacks them. The structure of the embedding, not just the vocabulary, shapes what is licensed. What does all this mean? Hacquard and Wellwood argue that the data force a two-part conclusion, and neither part alone is sufficient. First, epistemic modals are semantically contentful. The evidence from embedding proves it. If epistemics were purely speaker-level markers — like a hedging tone of voice — they could not appear in the scope of other operators in naturalistic text. But they do. That zero point ninety-two percent of "might" tokens appearing in complements of inquisitive verbs isn't noise; it's the language revealing that "might" can take scope inside a larger structure. Second, their distribution is still far more restricted than root modals. That restriction demands explanation, and it seems to arise from two distinct sources. In questions and conditional antecedents, the constraint appears pragmatic. Epistemic modals are anchored to a knowledge state — typically the speaker's. When you are asking a question or setting up an "if" clause, the pragmatic context often makes it infelicitous to invoke your own personal uncertainty. However, that bar can be cleared when the modal is anchored to a different knowledge state: an addressee, or a collective "we." Hacquard and Wellwood note a suggestive pattern — for "can," about three-quarters of inquisitive contexts are solipsistic, speaker-oriented. For "might," that bias disappears. Roughly half the inquisitive contexts involving "might" are addressee-oriented. When the anchoring shifts away from the speaker, epistemic "might" becomes more comfortable. In attitude contexts, the constraint appears more semantic. Representational attitude predicates — those that model an agent's information state, like "believe" or "think" — provide the kind of anchoring that epistemics require. They create an information state to which the modal can refer. Non-representational predicates — desideratives and directives — do not provide that information state, and thus do not license epistemics. Hacquard and Wellwood connect this directly to Anand and Hacquard's prior work on the split between representational and non-representational attitudes. The corpus data do not just confirm that split; they add texture. Emotive doxastic predicates like "hope" can license epistemic possibility modals like "might" but rarely license epistemic necessity like "must," consistent with the idea that "hope" contains a doxastic possibility component but not the certainty associated with "must." One open puzzle the authors flag is an asymmetry between "might" and "must." "Might" embeds more freely than "must" across environments, and that asymmetry is visible throughout the data but not yet fully explained. It's a thread the study leaves productively unresolved. What Hacquard and Wellwood have shown, ultimately, is that corpus evidence can break a theoretical deadlock that constructed examples could not. Real language, at scale, demonstrates that epistemic modals do embed — but that their embedding is licensed, not free. They are not illocutionary overlays floating above sentences; they are semantic operators that contribute content, but operators that require particular conditions to appear: pragmatic conditions surrounding knowledge anchoring in questions and conditionals, and semantic conditions tied to the representational structure of attitude predicates. The answer to whether "might" means something turns out to be yes — and also, it depends on who is doing the knowing. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

The word "might" in the sentence "John thinks it might rain" either means something or it doesn't. That sounds like a trivially simple question, but it isn't. For decades, linguists have split into two camps over exactly this issue, and the answer has consequences for how we think about meaning itself. One side argues that epistemic modals — words like "might," "must," and "can" when they express what is possible or necessary given what we know — are genuine contributors to truth conditions. They are part of what makes a sentence true or false. The other side claims they float above the sentence like a tone of voice, marking the speaker's attitude rather than shaping the proposition. The debate has been hard to settle because the evidence supports both sides. Hacquard and Wellwood decided to stop arguing with constructed examples and to examine how the language actually behaves. The theoretical stakes are worth considering for a moment. In the Kratzerian tradition — named for the semanticist Angelika Kratzer — epistemic modals are treated as quantifiers over possible worlds. The sentence "John must be the murderer" is true if, in all worlds compatible with what is known, John is the murderer. That's a real semantic contribution. The rival view treats epistemic expressions as speaker assessments — evidential markers that report the speaker's degree of commitment without changing the underlying proposition. The diagnostic that both camps have leaned on is embedding.

If epistemic modals can appear inside the scope of other operators — within questions, in the "if" clause of a conditional, or in the complement of a verb like "think" or "believe" — then they are behaving like ordinary semantic operators. If they cannot, that suggests they exist outside the sentence's meaning, at the speaker level. Classical accounts claimed that epistemics generally do not embed, but that claim turned out to be harder to defend than it appeared. Hacquard and Wellwood drew their data from the New York Times section of the English Gigaword Corpus, a massive written corpus that they parsed into fifteen million, six hundred ninety-one thousand, eight hundred fifty-nine sentences. Across that data, they found one hundred forty-nine thousand, two hundred nineteen tokens of "might," eighty-eight thousand, eight hundred fifty-nine of "must," and four hundred seventy-five thousand, five hundred ninety of "can." They examined three embedding environments: antecedents of conditionals, questions including embedded questions, and complements of attitude predicates. For "must" and "have to," they hand-annotated whether each token was epistemic or root, using paraphrases as guides — epistemic "must" reads as "it is probable that," while root "must" indicates an obligation. Interannotator agreement was high, with a kappa of 0.84. The parsed corpus allowed systematic comparison across environments at a scale that no constructed-example approach could match.

Now to what they actually found — and this is where the picture becomes genuinely interesting. Starting with conditional antecedents, the "if" clause in "If it might rain, bring an umbrella" turns out to be nearly hostile to epistemic modals in real data. Only thirty tokens of "might" occurred in antecedents of conditionals — that's 0.02 percent of all "might" tokens in the corpus. "Can," by contrast, appeared there nine thousand, two hundred ninety-two times, or one point ninety-five percent of its corpus distribution. For "must," there were two hundred thirteen total tokens in "if" clauses, and in all but one case, the modal received a root interpretation. Epistemic "must" in conditional antecedents was effectively absent. That's not a categorical ban from theory — it's an empirical near-zero from actual language use. Questions tell a more complicated story. In matrix questions — questions sitting at the top of a sentence — "can" is far more common than "might": three point seventy-eight percent of "can" tokens versus zero point thirty-five percent of "might" tokens appear in matrix questions. That asymmetry fits the speaker-oriented picture; asking a question about your own uncertainty feels strange.

However, embedded questions — questions embedded inside a larger sentence — do not show the same scarcity. "Might" appears in embedded questions at one point fifteen percent of its distribution, "can" at zero point eighty percent, and eighty-six cases of epistemic "must" are attested there. In complements of inquisitive verbs specifically, zero point ninety-two percent of "might" tokens appear, versus only zero point zero four percent of "can." So the ban is not on embedding in questions generally; it's more targeted than that. The richest evidence comes from attitude predicates — verbs like "think," "believe," "say," "want," and "order." This is where the data push back against a simple story. Epistemic modals do appear in the complements of some attitude verbs — abundantly, in fact, under verbs of acceptance, perception, certainty, and conjecture. However, under desideratives and directives — "want," "order," "demand" — they're essentially absent. The paper reports no instances of epistemic "have to" in desiderative or directive complements. Additionally, there's a revealing contrast within the attitude class: finite complements can host epistemic readings, while infinitival "have to" completely lacks them. The structure of the embedding, not just the vocabulary, shapes what is licensed. What does all this mean? Hacquard and Wellwood argue that the data force a two-part conclusion, and neither part alone is sufficient.

First, epistemic modals are semantically contentful. The evidence from embedding proves it. If epistemics were purely speaker-level markers — like a hedging tone of voice — they could not appear in the scope of other operators in naturalistic text. But they do. That zero point ninety-two percent of "might" tokens appearing in complements of inquisitive verbs isn't noise; it's the language revealing that "might" can take scope inside a larger structure. Second, their distribution is still far more restricted than root modals. That restriction demands explanation, and it seems to arise from two distinct sources. In questions and conditional antecedents, the constraint appears pragmatic. Epistemic modals are anchored to a knowledge state — typically the speaker's. When you are asking a question or setting up an "if" clause, the pragmatic context often makes it infelicitous to invoke your own personal uncertainty. However, that bar can be cleared when the modal is anchored to a different knowledge state: an addressee, or a collective "we." Hacquard and Wellwood note a suggestive pattern — for "can," about three-quarters of inquisitive contexts are solipsistic, speaker-oriented. For "might," that bias disappears. Roughly half the inquisitive contexts involving "might" are addressee-oriented. When the anchoring shifts away from the speaker, epistemic "might" becomes more comfortable.

In attitude contexts, the constraint appears more semantic. Representational attitude predicates — those that model an agent's information state, like "believe" or "think" — provide the kind of anchoring that epistemics require. They create an information state to which the modal can refer. Non-representational predicates — desideratives and directives — do not provide that information state, and thus do not license epistemics. Hacquard and Wellwood connect this directly to Anand and Hacquard's prior work on the split between representational and non-representational attitudes. The corpus data do not just confirm that split; they add texture. Emotive doxastic predicates like "hope" can license epistemic possibility modals like "might" but rarely license epistemic necessity like "must," consistent with the idea that "hope" contains a doxastic possibility component but not the certainty associated with "must." One open puzzle the authors flag is an asymmetry between "might" and "must." "Might" embeds more freely than "must" across environments, and that asymmetry is visible throughout the data but not yet fully explained. It's a thread the study leaves productively unresolved.

What Hacquard and Wellwood have shown, ultimately, is that corpus evidence can break a theoretical deadlock that constructed examples could not. Real language, at scale, demonstrates that epistemic modals do embed — but that their embedding is licensed, not free. They are not illocutionary overlays floating above sentences; they are semantic operators that contribute content, but operators that require particular conditions to appear: pragmatic conditions surrounding knowledge anchoring in questions and conditionals, and semantic conditions tied to the representational structure of attitude predicates. The answer to whether "might" means something turns out to be yes — and also, it depends on who is doing the knowing. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Arts and Humanities