The Effect of Customers' Emotional Responses to Service Failures on Their Recovery Effort Evaluations and Satisfaction Judgments
Picture the moment a service craters on you. The hotel room isn't ready after a long flight. Your entree arrives cold.
In that instant, you don't pull out a spreadsheet of expectations and justice theories. You feel something first. The big question is whether those feelings just color the moment or actually change how you judge the company's recovery, how you weigh the apology, the speed, and the compensation.
Smith and Bolton set out to answer that in a way that goes beyond "people get mad" and into what that anger does inside the evaluation machinery.
They designed a controlled, two-industry test—restaurants and hotels—to separate the cognitive levers we usually talk about from the raw emotional response. A service failure, in their setup, either hit the outcome itself or the way it was delivered, and it came in two sizes: high magnitude or low. Then came recovery with familiar ingredients: compensation that was none, medium, or high; a response that was immediate or delayed; an apology present or absent; and whether the organization took the initiative or waited for the customer to flag the issue.
On the back end, they measured two kinds of satisfaction. One is transaction-specific—how satisfied you are with that failure and recovery encounter. The other is cumulative—your overall satisfaction with the firm across experiences.
That split matters because a blowout tonight might be a blip in the long run. Or it might not.
Here's the clever part. Instead of asking people how they felt after they read the recovery, they captured emotion right after the failure and before any recovery information could muddy the waters. Participants did a brief think aloud—just their thoughts and feelings in the moment.
The researchers later coded these for five negative emotions drawn from Oliver and Richins: anger, discontent, disappointment, self-pity, and anxiety. Then everyone read one of the recovery scenarios and evaluated it. That sequencing lets us ask: do negative feelings at time one change how you process the recovery at time two?
The samples give us both breadth and control. In restaurants, three hundred fifty-five undergraduates each described a recent non-fast-food visit. In hotels, five hundred forty-nine business travelers from a midrange chain's reservation list completed a mail survey with a cash incentive.
The undergraduates bring variety across many brands; the hotel sample is tighter, anchored to a single chain. And across both, the emotional split is striking. In the restaurant study, fifty-nine percent voiced at least one negative emotion in that verbal snapshot.
In hotels, it was forty percent. Overall, forty-seven percent of all respondents showed emotion, while fifty-three percent did not. That's not a tiny corner case. It's half the market.
Now, how do you model something as squishy as "feeling bad"? Smith and Bolton did two things. First, they ran parallel equations for the emotion group and the no-emotion group.
Same predictors—expectations, perceived performance, the sense of disconfirmation (that gap between what you expected and what you got), and three flavors of justice: distributive (outcomes), procedural (process fairness), and interactional (how respectfully you were treated). If the intercept for the emotion group is lower, that's a direct penalty of feeling negative. If the slopes differ—if justice matters more or less when you're emotional—that's a moderation effect.
Second, they zoomed in on recovery mechanics. In a recovery-focused model, satisfaction depends on failure type and magnitude and on the recovery knobs—initiation, apology, speed, and compensation. It also depends on those same cognitive judgments. Think of it as "what they did" and "how you interpreted it," split cleanly.
There's one more elegant tweak under the hood. People who felt something also tended to think more. Across the full sample, emotional respondents generated about three thoughts on average, versus roughly two and a half for those without emotion.
That difference isn't just a curiosity; it creates noise. So the researchers estimated their models with weighted least squares, giving less weight to respondents with long thought chains. Formally, the weight is the inverse of the number of thoughts.
That way, a chatty transcript doesn't overwhelm the regression. They checked error variances with a standard Glesjer test. Errors were equal across the restaurant equations and different in the hotel equations, which justifies estimating hotels separately.
It's a reminder that the statistics have to bend to the psychology, not the other way around.
So what did they find? Start with the headline: in hotels, emotion reshapes the evaluation of recovery. In restaurants, not so much.
A set of Chow tests tells the story. When they asked whether the emotion and no-emotion equations could be pooled together in restaurants, the answer was yes; the F statistic wasn't even close to rejection. In hotels, pooling was rejected.
In the combined dataset—restaurants and hotels together—it was rejected decisively. Translation: in hotels and in the overall view, customers who felt negative emotion didn't just start from a worse mood. They used the recovery information differently.
How much worse is "worse"? In the hotel equations, the intercept for the emotion group sits roughly two-thirds of a point lower on a seven-point satisfaction scale than the no-emotion group. That's the main effect: same recovery, same justice perceptions, but starting lower if you felt bad at the failure.
And then moderation kicks in. The coefficients on performance, disconfirmation, and the three justices differ significantly between emotion and no-emotion in hotels. In plain terms, when a customer is emotional, they tune the dials on what matters. Outcome fairness looms larger. Politeness, by contrast, fades.
The variance story makes this concrete. Among hotel guests who felt negative emotion, almost all the explained variation in encounter satisfaction was tied to the levers of recovery performance and to those cognitive antecedents. About ninety-eight percent of it.
In the no-emotion hotel group, the same bundle explained roughly three-quarters. That twenty-four point gap means that recovery attributes and fairness perceptions pay off more when customers are emotional. It also lines up with how the weights shift.
In the hotel data, distributive justice—the sense that the eventual outcome was fair—dominates in the emotion group. Interactional justice—the courtesy and explanations—plays a much bigger role for the no-emotion group. If you're upset, make the outcome right. If you're not, how you're treated can carry the day.
Restaurants are the foil here. Despite a higher share of emotional transcripts, emotion didn't exert a reliable main effect on encounter satisfaction in that setting. Nor did it change the weights on performance, disconfirmation, or justice.
It's not that recovery didn't matter; adjusted R-squared values were strong. But the presence of negative emotion didn't systematically rewire the evaluation. That difference across settings is part psychology, part context.
A hotel stay stacks multiple touchpoints and higher stakes; a restaurant meal is shorter and, often, cheaper. The data reflect that.
Now, can an excellent recovery overcome the emotional penalty? The pooled analysis offers a neat nuance. When you combine industries and imagine a "typical" recovery, the negative main effect of emotion is almost perfectly offset.
It is offset by its positive interaction with the recovery attributes. The numbers practically cancel: a hit just under one point on the main effect, a boost just under one point on the interaction. Net, near zero.
Push the recovery from typical to excellent, and the interaction wins. In other words, if someone's upset, there's more room to help yourself—and more room to hurt yourself—through what you do next.
Let's ground the stakes with a few more facts from the design. The manipulation of failure severity worked. People consistently rated the high-magnitude failures as more severe than the low-magnitude ones across both services.
Severity didn't differ by failure type. Emotion categories showed the heterogeneity you'd expect. Restaurants had roughly double the incidence of anger compared with hotels, and discontent was common in both.
But these discrete labels were used for coding, not modeled separately. The average number of thoughts was higher among those who were emotional. That supports the idea that feeling bad goes hand in hand with more systematic processing rather than mindless rage.
Underneath the hood, model fit was solid. In hotels, the encounter-satisfaction equations explained more than four-fifths of the variance for both emotion and no-emotion groups. In restaurants, fit was somewhat lower but still strong.
And when the authors looked at overall, cumulative satisfaction, emotion mattered less than it did for the specific encounter. The gap is small in absolute terms but meaningful: emotion explained more variance in the transaction than in the cumulative measure. That aligns with intuition.
Feelings at time one press down hard on judgments at time two, then dissipate as other experiences pile on.
What do you do with all this if you run a front desk or a dining room? Start by reading the room. Smith and Bolton's coding list—anger, discontent, disappointment, self-pity, anxiety—maps closely to what you can see and hear.
When those signals are present in hotels, the recovery levers are more potent. That means speed and compensation aren't just nice to have; they're the tools that shift satisfaction when it counts. It also means that fairness in outcomes—what the customer ultimately gets—beats a graceful apology when someone is upset.
Courtesy still matters, especially for customers who aren't already in a negative state. But it won't save you if the outcome stays wrong.
A brief word on limitations keeps us honest. These were hospitality contexts with one midrange hotel chain on the lodging side and a wide mix of restaurants on the other. That structure strengthens internal validity in hotels and external validity in restaurants, but it also makes pooling tricky.
The emotion measure came from verbal protocols, which is great for avoiding halo effects. But it doesn't directly translate to settings where emotion is hard to elicit or observe. And while the team cataloged five discrete negative emotions, they modeled an overall "negative emotion" construct because category sizes were small.
Intensity, interestingly, didn't add explanatory power and was dropped from the final models.
The deeper takeaway is conceptual. Emotions aren't just noise on the satisfaction channel; they can retune the channel. In hotels, negative feeling shifts the weight onto concrete outcomes and amplifies the value of doing recovery well.
In restaurants, that shift doesn't show up in the same way, which hints at boundary conditions. Service length, stakes, and the number of touchpoints likely matter. It's also a caution against one-size-fits-all advice.
A Chow test is not a crescendo in a podcast, but it's a sober reminder. Combining subsamples is not always appropriate, and the implications can differ by setting.
If you're thinking about what to study next, two paths are obvious. One is generalization. Retail, airlines, health care—do we see the same moderation pattern?
The other is precision. Instead of lumping negative feelings together, isolate anger from anxiety, disappointment from discontent, and ask whether each one tilts the evaluation in its own way. Smith and Bolton opened that door with careful coding and clean sequencing. The next studies can walk through it.
But you don't need the next study to act on the core lesson. When a failure happens, listen for emotion first. If it's there, dial up the outcome fixes—speed, compensation, making it right—and know that those moves will punch above their weight.
If it's not, lean into the experience—the explanation, the tone, the respect—because that's what will carry the day. Either way, you're not just managing a recovery. You're managing the way people think while they feel.
Related lectures
- Responding to the ripple effect from systemic disruptions: empirical evidence from the semiconductor shortage during COVID-19
- A whole new world: Counterintuitive crowdfunding insights for female founders
- Unveiling founder archetypes: the effect of distinct entrepreneurial traits on resource accuracy
- Burnout and engagement at work as a function of demands and control
- Greenwashing and environmental communication: Effects on stakeholders' perceptions
- Improved Response to Disasters and Outbreaks by Tracking Population Movements with Mobile Phone Network Data: A Post-Earthquake Geospatial Study in Haiti