“I don’t Think These Devices are Very Culturally Sensitive.”—Impact of Automated Speech Recognition Errors on African Americans

Zion Mengesha, Courtney Heldreth, Michal Lahav, Juliana Sublewski, Elyse TuennermanView original
OverviewBalancedadam voice
Someone picks up their phone and tries to send a voice message. They get halfway through a sentence, stop, and start again — this time slower, flatter, and stripped of the rhythm that normally runs through their speech. They didn't decide to do that consciously. It just happened because it has to happen if they want the thing to work. Mengesha and colleagues at Google set out to understand exactly what that moment costs — not in error rates, but in the person doing the accommodating. The technical disparity was already documented before this paper. Automated speech recognition, or ASR, which is the software that converts spoken language into text and powers virtual assistants, dictation tools, and hands-free computing, performs measurably worse for many speakers of African American Vernacular English, or AAVE, a systematic, rule-governed variety of English used in African American communities. Koenecke and colleagues, cited by Mengesha and colleagues, found that African American speakers can face word error rates up to two times higher than white speakers of standard American English. Two times. That's not a rounding error in a benchmark — that's the difference between a tool that works and one that doesn't. However, the research before this paper mostly stopped at measuring the gap. What nobody had carefully documented was the downstream consequence: what it actually does to a person when technology consistently mishears them. Does it frustrate them? Embarrass them? Change how they think about themselves? Mengesha and colleagues built a study to find out. They chose a diary study design, and the choice matters. Lab experiments capture what people do when they know they're being watched; surveys capture what people remember and can articulate on demand. A diary study captures what actually happens in daily life. Participants were recruited through the dScout mobile platform — a screening of one thousand eight hundred sixty-five people filtered down to thirty African American, native English speakers who used voice technology regularly and had experienced errors they believed were linked to how they speak. Participants were drawn from Atlanta, Chicago, Houston, Los Angeles, New Orleans, Philadelphia, and Washington D.C., paid one hundred fifty dollars each, and sampled to balance age, gender, income, and education. The protocol ran for two weeks and had five parts. The first day was a baseline survey. Then followed five days of in-context diary logging: every time a participant used voice technology, they reported what they were doing, what happened, and submitted a sixty-second video. After that, two days focused specifically on negative experiences — open descriptions, closed ratings, and screen recorded recreations of failures. The final day assigned participants to one of seven specific voice tasks: sending a message, setting a reminder, writing an email, running a search, getting directions, or calling a contact. By the end, the team had six hundred forty unique open-ended responses, one thousand eighty closed-ended responses, and two hundred forty transcribed videos, analyzed through inductive, bottom-up thematic coding that ultimately produced one hundred twenty-four individual codes. That is a lot of data grounded in real, naturalistic behavior. And what it shows is not merely that ASR fails these users — it's what those failures mean to them. The dominant finding is othering. Participants did not experience ASR errors as random technical glitches. They experienced them as signals. When a system misheard them, it prompted thoughts about race and place — about whether they belonged to the category of people the technology was built to serve. One participant, identified as P7 from Chicago, put it plainly: "The technology is made for the standard middle-aged white American, which I am not." That's not frustration talking. That's a conclusion drawn from repeated experience. The numbers behind that conclusion are striking. When asked why voice technology fails them, thirty percent said the technology wasn't designed to pick up accents and slang, twenty percent pointed to their speech patterns, and ten percent cited names or vocabulary. When asked who the technology works better for, thirty-six percent said white people, and another thirty-six percent said people without an accent. The emotional toll was substantial: seventy-seven percent reported frustration during failures, fifty-eight percent felt bothered, fifty-five percent felt disappointed, and fifty-two percent felt angry. Thirty-six percent reported anxiety. These aren't annoyances. They're the emotional signature of being consistently excluded. And the errors had real-world consequences. Thirty-six percent of participants reported dissatisfaction when mis-transcription produced wrong results, thirty-two percent when the system didn't understand their commands, and thirty-two percent when they had to complete the task manually anyway. One participant described a mis-transcription that "conveyed the opposite message than what I had originally intended, and cost somebody else a lot of time." Now add the behavioral layer. Ninety-three percent of participants — twenty-eight out of thirty — reported modifying their dialect to be understood by voice technology. Nearly everyone. They spoke slower, avoided slang, and clipped the cadences out of their sentences. One participant described it as "talking real clear, and don't use slang words like my regular talk." Prior work cited by Mengesha and colleagues shows speakers even lengthen vowel durations after an ASR error — fine-grained phonetic adjustments happening in real time, below the level of conscious decision. Linguists call this accommodation — deliberately modifying your speech to match an expected standard. In some contexts, it's a social skill. Here, it's a tax. Of the twenty-eight participants who said they changed their speech, sixty-seven percent felt bothered by having to do so, fifty-three percent felt frustrated, forty percent felt disappointed, thirty-three percent felt angry, and seventeen percent felt self-conscious. One participant made the stakes explicit: "It needs to change because it doesn't feel inclusive when I have to change how I speak and who I am, just to talk to technology." Another captured the practical futility: "I might as well have typed it out myself instead of going back again rereading every word, deleting words, and adding words." What makes the accommodation finding especially pointed is where the burden lands. A majority of participants — fifty-four percent — agreed or strongly agreed that they needed to modify their speech because the technology doesn't understand their racial group. And yet seventeen percent blamed themselves for speaking too fast. The system has successfully redistributed its own failure onto its users. That's not a design accident. It's what happens when the people building a tool don't include the full range of people who will use it. Mengesha and colleagues close with four directions, each tied to a specific failure mode the study identified. First: diversify training data. AAVE and other underrepresented varieties need to be in the datasets ASR systems learn from — across region, gender, age, and socioeconomic status — with careful attention to privacy and consent. Participants said they were willing to contribute voice samples; the paper flags that willingness should be met with transparency. Second: develop personalized speech models, giving users correction pathways and what the authors call federated repair — adaptive systems that learn from individual users. Third: investigate dialectical transcription preferences, meaning how African American users actually want their speech transcribed, not just whether the transcription is technically accurate. Fourth: involve community voices through community-based participatory research from the beginning of the design process, not as an afterthought. And then there's the methodological argument, which is worth naming on its own. Mengesha and colleagues make a case that the diary study is a tool for AI fairness research — not just for this question but for any population whose experiences with technology are salient, occasional, and shaped by identity. Lab studies and surveys miss the texture of daily life. A two-week diary captures the pattern across many moments, in context, reported by the person living it. The deeper stakes here are not really about voice assistants. They're about legibility — about who gets to be understood by the systems that increasingly mediate everyday life. When ninety-three percent of a group reports having to change how they speak to be heard by a machine, the machine has a problem. The question this paper leaves open is whether the people building those machines will treat it like one. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Someone picks up their phone and tries to send a voice message. They get halfway through a sentence, stop, and start again — this time slower, flatter, and stripped of the rhythm that normally runs through their speech. They didn't decide to do that consciously. It just happened because it has to happen if they want the thing to work. Mengesha and colleagues at Google set out to understand exactly what that moment costs — not in error rates, but in the person doing the accommodating. The technical disparity was already documented before this paper. Automated speech recognition, or ASR, which is the software that converts spoken language into text and powers virtual assistants, dictation tools, and hands-free computing, performs measurably worse for many speakers of African American Vernacular English, or AAVE, a systematic, rule-governed variety of English used in African American communities. Koenecke and colleagues, cited by Mengesha and colleagues, found that African American speakers can face word error rates up to two times higher than white speakers of standard American English. Two times. That's not a rounding error in a benchmark — that's the difference between a tool that works and one that doesn't. However, the research before this paper mostly stopped at measuring the gap. What nobody had carefully documented was the downstream consequence: what it actually does to a person when technology consistently mishears them. Does it frustrate them?

Embarrass them? Change how they think about themselves? Mengesha and colleagues built a study to find out. They chose a diary study design, and the choice matters. Lab experiments capture what people do when they know they're being watched; surveys capture what people remember and can articulate on demand. A diary study captures what actually happens in daily life. Participants were recruited through the dScout mobile platform — a screening of one thousand eight hundred sixty-five people filtered down to thirty African American, native English speakers who used voice technology regularly and had experienced errors they believed were linked to how they speak. Participants were drawn from Atlanta, Chicago, Houston, Los Angeles, New Orleans, Philadelphia, and Washington D.C., paid one hundred fifty dollars each, and sampled to balance age, gender, income, and education. The protocol ran for two weeks and had five parts. The first day was a baseline survey. Then followed five days of in-context diary logging: every time a participant used voice technology, they reported what they were doing, what happened, and submitted a sixty-second video.

After that, two days focused specifically on negative experiences — open descriptions, closed ratings, and screen recorded recreations of failures. The final day assigned participants to one of seven specific voice tasks: sending a message, setting a reminder, writing an email, running a search, getting directions, or calling a contact. By the end, the team had six hundred forty unique open-ended responses, one thousand eighty closed-ended responses, and two hundred forty transcribed videos, analyzed through inductive, bottom-up thematic coding that ultimately produced one hundred twenty-four individual codes. That is a lot of data grounded in real, naturalistic behavior. And what it shows is not merely that ASR fails these users — it's what those failures mean to them. The dominant finding is othering. Participants did not experience ASR errors as random technical glitches. They experienced them as signals. When a system misheard them, it prompted thoughts about race and place — about whether they belonged to the category of people the technology was built to serve. One participant, identified as P7 from Chicago, put it plainly: "The technology is made for the standard middle-aged white American, which I am not." That's not frustration talking. That's a conclusion drawn from repeated experience.

The numbers behind that conclusion are striking. When asked why voice technology fails them, thirty percent said the technology wasn't designed to pick up accents and slang, twenty percent pointed to their speech patterns, and ten percent cited names or vocabulary. When asked who the technology works better for, thirty-six percent said white people, and another thirty-six percent said people without an accent. The emotional toll was substantial: seventy-seven percent reported frustration during failures, fifty-eight percent felt bothered, fifty-five percent felt disappointed, and fifty-two percent felt angry. Thirty-six percent reported anxiety. These aren't annoyances. They're the emotional signature of being consistently excluded. And the errors had real-world consequences. Thirty-six percent of participants reported dissatisfaction when mis-transcription produced wrong results, thirty-two percent when the system didn't understand their commands, and thirty-two percent when they had to complete the task manually anyway. One participant described a mis-transcription that "conveyed the opposite message than what I had originally intended, and cost somebody else a lot of time." Now add the behavioral layer. Ninety-three percent of participants — twenty-eight out of thirty — reported modifying their dialect to be understood by voice technology. Nearly everyone.

They spoke slower, avoided slang, and clipped the cadences out of their sentences. One participant described it as "talking real clear, and don't use slang words like my regular talk." Prior work cited by Mengesha and colleagues shows speakers even lengthen vowel durations after an ASR error — fine-grained phonetic adjustments happening in real time, below the level of conscious decision. Linguists call this accommodation — deliberately modifying your speech to match an expected standard. In some contexts, it's a social skill. Here, it's a tax. Of the twenty-eight participants who said they changed their speech, sixty-seven percent felt bothered by having to do so, fifty-three percent felt frustrated, forty percent felt disappointed, thirty-three percent felt angry, and seventeen percent felt self-conscious. One participant made the stakes explicit: "It needs to change because it doesn't feel inclusive when I have to change how I speak and who I am, just to talk to technology." Another captured the practical futility: "I might as well have typed it out myself instead of going back again rereading every word, deleting words, and adding words." What makes the accommodation finding especially pointed is where the burden lands. A majority of participants — fifty-four percent — agreed or strongly agreed that they needed to modify their speech because the technology doesn't understand their racial group. And yet seventeen percent blamed themselves for speaking too fast.

The system has successfully redistributed its own failure onto its users. That's not a design accident. It's what happens when the people building a tool don't include the full range of people who will use it. Mengesha and colleagues close with four directions, each tied to a specific failure mode the study identified. First: diversify training data. AAVE and other underrepresented varieties need to be in the datasets ASR systems learn from — across region, gender, age, and socioeconomic status — with careful attention to privacy and consent. Participants said they were willing to contribute voice samples; the paper flags that willingness should be met with transparency. Second: develop personalized speech models, giving users correction pathways and what the authors call federated repair — adaptive systems that learn from individual users. Third: investigate dialectical transcription preferences, meaning how African American users actually want their speech transcribed, not just whether the transcription is technically accurate. Fourth: involve community voices through community-based participatory research from the beginning of the design process, not as an afterthought.

And then there's the methodological argument, which is worth naming on its own. Mengesha and colleagues make a case that the diary study is a tool for AI fairness research — not just for this question but for any population whose experiences with technology are salient, occasional, and shaped by identity. Lab studies and surveys miss the texture of daily life. A two-week diary captures the pattern across many moments, in context, reported by the person living it. The deeper stakes here are not really about voice assistants. They're about legibility — about who gets to be understood by the systems that increasingly mediate everyday life. When ninety-three percent of a group reports having to change how they speak to be heard by a machine, the machine has a problem. The question this paper leaves open is whether the people building those machines will treat it like one. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Arts and Humanities