Exploring the impact of word-of-mouth about Physicians’ service quality on patient choice based on online health communities
For most of human history, choosing a doctor meant asking a neighbor, a relative, or no one at all. Then, a single Chinese website changed that. One thousand eight hundred and fifty-three physicians' real booking records allowed researchers to finally measure what patients actually respond to when they pick a doctor online. Not what patients say they care about, but what they do. That distinction turns out to matter enormously because the answer is more complicated than anyone designing those rating systems assumed. Health care is what economists call a high-credence service. Naiji Lu and Hong Wu define the category precisely as services so saturated with uncertainty that patients often cannot evaluate quality even after receiving care. A patient can tell whether a restaurant meal tasted good, but they usually cannot tell whether a diagnosis was accurate, whether a treatment was optimal, or whether a different physician would have done better. Judgments depend heavily on the provider's ability to communicate. That opacity raises the stakes for any information signal a patient can actually see, and it's exactly why online ratings have become so consequential.
Before platforms like Guahao.com and Haodf.com existed, most research on health care quality focused on institutions such as hospitals or nursing homes and not individual physicians. There was simply no public data on individual doctors at scale. Online health communities changed that by introducing word-of-mouth review functions where patients share ratings and written experiences after each visit. For prospective patients, those ratings are now a primary tool for navigating an otherwise opaque market. The question Lu and Wu set out to answer was: when patients can actually see what other patients said, what drives the booking decision? To answer it, they needed a framework for what "quality" means in this context. They use a two-dimensional structure that traces back to health services researcher Donabedian. Technical quality is the outcome of care, meaning did the treatment work? Functional quality is the process — how did the physician treat you as a person? Bedside manner, communication, and attentiveness. On Guahao.com, these map directly onto the two rating categories patients fill out: treatment outcome and physician attitude. The distinction isn't just academic. Technical and functional quality speak to fundamentally different patient needs. Prior work treats technical quality as the core driver of satisfaction, as the clinical result is what you came for.
Functional quality, by contrast, influences emotional well-being and the patient's perception of competence. Research cited by Lu and Wu shows that socio-emotional behaviors such as listening, caring, and taking time are central to patients' sense of being cared for. But here's the non-obvious part: because they meet different needs, these two dimensions need not simply add up. High marks on one might compensate for lower marks on the other, or they might amplify each other. Lu and Wu formalize both possibilities as competing hypotheses and then let the data decide. The data come from Guahao.com, China's largest online health platform, with over twenty-three million registered users and roughly one hundred thousand listed physicians. Lu and Wu collected profiles and patient feedback for one thousand eight hundred and fifty-three physicians across sixteen diseases: eight high-risk malignant tumor conditions, with a mortality rate of about one hundred sixty-four deaths per one hundred thousand people in 2013, and eight lower-risk endocrine-related conditions, with mortality around seventeen per one hundred thousand. That roughly tenfold difference in stakes is the lever they use to test whether disease severity shifts how patients weigh quality signals.
The outcome variable is the change in online bookings between two data collection points one month apart — the log of a physician's total appointments at the second time minus the first time. Using earlier-period ratings to predict later booking changes helps establish that quality signals precede the patient choices they're associated with. The estimation method is ordinary least squares regression, which identifies which factors statistically predict that change in bookings while holding everything else constant. Controls include hospital tier, hospital environment scores, physician title, and logged experience. Crucially, the models also include an interaction term between technical and functional quality, along with interaction terms between each quality dimension and the high-risk disease dummy. The headline result is the one you wouldn't have predicted. Functional quality — bedside manner and attitude — negatively moderates the relationship between technical quality and patient choice. The interaction coefficient between treatment outcome and attitude is negative two point nine seven zero, significant at a p-value below zero point zero zero one.
That negative sign means that when a physician already scores high on warmth and communication, the additional booking boost from excellent treatment outcomes is reduced. The two dimensions substitute for each other at least at the margin. Good bedside manner can partially replace strong clinical outcome signals in the patient's decision calculus. Both dimensions still matter on their own. The main effect of treatment outcome is zero point seven five eight, and the main effect of attitude is zero point eight eight nine, both significant at a p-value below zero point zero zero one. Higher scores on either dimension are associated with more future bookings. But they don't simply add up — the negative interaction means the combined effect is less than the sum of the parts. Disease risk then reshapes the whole picture. For high-risk diseases, the interaction between treatment outcome and the risk dummy is one point one five five, indicating that patients weigh clinical outcomes more heavily when the stakes are high. For low-risk diseases, the attitude coefficient is larger, and the disease-risk interaction with attitude is negative, meaning functional quality carries relatively more weight when the condition is less severe. The robustness checks, which involve splitting the sample into high-risk and low-risk subgroups and rerunning the models, confirm the substitution pattern persists in both. The full model explains eighty-one percent of the variance in booking changes.
Taken together, the pattern implies something specific about how patients process information under uncertainty. They don't appear to maximize a weighted sum of all available quality signals. Instead, they satisfice — they accept a combination that clears some threshold, using one strong signal to compensate for a weaker one. When bedside manner is already high, excellent outcomes are less marginal; the patient's threshold is already met. When the disease is serious enough that being wrong could be catastrophic, patients tighten that rule and tilt hard toward outcome signals. The fairness-heuristic logic Lu and Wu invoke captures this well: a pleasant process compensates for a less-than-perfect outcome, but only when the cost of the outcome falling short is survivable. For physicians and the managers who oversee them, the implications are practical and specific. Improving technical skill alone or bedside manner alone won't fully drive patient choice; both dimensions matter, and they interact. For physicians treating high-risk patients, investments in clinical outcomes generate larger booking returns. For those treating lower-risk conditions, warmth and attentiveness carry more relative weight. Because functional quality can partly substitute for technical quality, cultivating compassionate and respectful interactions is not merely being pleasant — it provides a partial hedge against the uncertainty inherent in any clinical outcome.
For online health platforms, the findings suggest that presenting quality as a single aggregate score, or even two independent scores, may not reflect how patients actually use that information. A platform that knows a patient is searching for an oncologist might reasonably surface technical outcome ratings more prominently than it would for a patient searching for a routine endocrinology appointment. Lu and Wu are candid about the limits of what this study can establish. The disease selection is coarse, the mortality difference between the two groups is roughly tenfold, and the analysis focuses on a single platform and service context. They recommend more granular disease classification, validation across other health care settings, and longitudinal designs that can better track how patient preferences evolve. But the core finding stands on its own. The next time a patient scrolls through physician ratings on a health platform, they are running a calculation that turns out to be contextual, interactive, and sensitive to the severity of what they're facing. The rating system shows them a number. What they do with it depends on which dimension that number captures, what the other dimension shows, and how much is at stake if they get the choice wrong. Lu and Wu are the first to measure all three of those variables at once, with real data, real patients, and real choices. This lecture was created by ennepō.
Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Determinants of the Pace of Global Innovation in Energy Technologies
- Brand negativity: a relational perspective on anti-brand community participation
- Quantitative ESG disclosure and divergence of ESG ratings
- Inferring Labor Income Risk and Partial Insurance From Economic Choices
- The stock market's reaction to quality certification: Empirical evidence from Spain
- Responding to Globalization: Impacts of Certification on Colombian Small-Scale Coffee Growers