Word-of-Mouth Communication and Social Learning

Glenn Ellison, Drew FudenbergView original
OverviewBalancedalloy voice
Here’s the puzzle Glenn Ellison and Drew Fudenberg put on the table: if people mostly learn by chatting with a few peers, when does that chatter make everyone pile onto the same choice, and when does it keep a mix alive? The twist that makes their 1995 paper so memorable is this: sometimes talking less pushes people to conform, and under the right conditions, that very conformity steers the whole population to the better action. Conformity isn’t always a bug; it can be a feature. They build the story in a stripped-down world with two technologies, which they call f and g. Each period, payoffs bounce around because of two kinds of shocks. There’s a common shock that hits everyone in the same direction, and there are idiosyncratic shocks that differ from person to person. Think weather versus whim. The common shock has two possible values, and there’s a probability p that the state favors g in any given period. The population's state is just one number, x sub t, the fraction currently using g. People don’t constantly rethink. A fixed share Q of the population reevaluates each period, and when they do, they look out at N randomly chosen peers, compare the average payoff they saw from f and from g in that sample, and switch to whichever had the higher sample average. It’s a very human rule: you look around, see how things went for a handful of others, and copy the better-looking option. It’s also not fully Bayesian. Because you only switch if you happen to talk to someone on the other side, popular options get observed more, which builds in a subtle popularity weighting. Under the hood, that simple sampling rule turns the whole system into a Markov process for x sub t. Given x sub t, N, Q, and the common shock for that period, you can write down the probability that a reevaluating person flips to g. In aggregate, that gives you the expected value of x sub t plus one. If you let N get large, the random sample shares look normal, and the update can be expressed with the standard normal cumulative distribution function—a smooth map that carries today’s share into tomorrow’s. But you don’t need the equation to follow the punchline. The dial that matters most is N, the number of people each person samples. Turn it down, and one kind of world emerges. Turn it up, and you get a different one. Ellison and Fudenberg start by looking at what happens near the edges—when a technology is so rare it’s barely hanging on. They prove a clean local result: you can linearize the dynamics around an endpoint and read off the drift. If the expected growth factor near zero is greater than one, then g can’t die out from that neighborhood; the process is repelled from zero. If the factor is less than one, there’s positive probability you slide all the way down to zero. The trick that makes this work is to analyze the logarithm of the market share near the boundary, where tiny changes in share correspond to big changes in the log. It sounds technical, and it is, but the intuition is simple—ask whether a vanishingly small minority is expected to grow or shrink on its next step. That local criterion ends up anchoring the global picture. With that tool in hand, they draw the first of their big dividing lines. For each strength of the common shock—call it omega—there’s a critical sample size N star of omega. Below it, the world tends to conformity: one of the two technologies becomes universal with probability one. Above it, the process does not settle on an endpoint. Shares keep moving, and diversity persists. That N star of omega curve rises with omega. Stronger common shocks push the boundary up; to get lock-in when the shared environment is volatile, you have to keep samples small. There’s a striking limiting case here that helps build intuition. If you turn off common shocks entirely and leave only idiosyncratic noise, the dynamics become deterministic and collapse to a simple split: for any N of at least two, the market share converges to one-half. There’s no conformity when weather-like forces disappear and inertia is positive. Now tilt the playing field so one option is better on average. Suppose p exceeds one-half, so in a typical period g is more likely to win. This is where their analysis sharpens because the model doesn’t just say “the better option wins.” It carves the parameter space into three zones using a function they call Q star of omega and N that comes out of the linearization. You compare that function to two simple ratios: p over one minus p and its reciprocal. If the ratio p divided by one minus p sits below Q star of omega and N, the population still locks in, but from some starting points it can lock onto the inferior technology—inefficient herding. If Q star of omega and N is sandwiched between those two ratios, everyone ends up on g—efficient social learning. If Q star of omega and N falls below one minus p divided by p, the system doesn’t converge to an endpoint at all; it stays diverse. It’s an elegant way to say that the average superiority of g helps, but only if the sampling rule lies in the right band. That band has structure. There isn’t just one boundary; there are two. The lower one is the N star of omega we just met, which marks the pivot from conformity to something else. There’s also an upper boundary, which we’ll call N prime of omega. When p is over one-half, the region in between—N star of omega to N prime of omega—is where learning is efficient and the population reliably finds and sticks with g. As p rises, that “Goldilocks” band widens. The lower boundary drops, and the upper boundary rises, meaning a broader swath of sample sizes deliver the right long-run outcome. There’s a tough-love message built in: outside that band, either herding can be wrong, or talking to too many people can keep the system from settling even when one option is better. You can attach some concrete numbers to these thresholds to get a feel for scale. They look at n, a way to summarize how strong the common shock is relative to idiosyncratic noise. When the common shock is twice as large as individual noise (n equals 2), the critical sample size that matters in their illustrations is on the order of forty. When the common shock is about the same size as individual noise (n equals 1), the comparable threshold drops to around five. The absolute figures depend on parameterization, but the lesson travels: in volatile shared environments, “ask fewer people” becomes a surprisingly good rule if you care about everyone eventually picking the better technology. Push N very high, and you can get stuck in motion—diversity that never collapses to a single winner and, in their terms, a process that doesn’t even enter a neighborhood of full convergence. So far, the conversation rule gave each observed payoff equal weight within the sample. Ellison and Fudenberg then ask, what if social influence leans harder on what’s popular? They add a parameter m that tilts sampling probabilities toward whatever is more prevalent. Think of it as a knob on the popularity amplifier. The three-regime map survives—conformity, efficient learning, diversity—but the borders slide. As Theorem 3 lays out, ratcheting up m generally makes convergence to an endpoint easier. In some initial conditions, that means the system can get pulled into inefficient lock-in that wouldn’t happen without the extra popularity weighting. The “best” sample size also shifts. The uniformly optimal rule—which in the baseline is N star of omega, the size that guarantees efficiency across all p when p is over one-half—becomes N star of omega and m and rises with m. In a numerical illustration the authors discuss, with a moderate common shock and p at 0.6, the sweet spot is just a handful of conversations—around three. That detail is less important than the principle: amplify popularity and you don’t need to sample as widely to herd, for better or worse. Then comes memory. Up to now, people were reacting only to what they and their peers just experienced. Add social memory—either a short, two-period lookback or, in the limit, an infinite one—and the role of the common shock changes. With even a bit of memory, fluky aggregate conditions get averaged out. That makes it easier for moderate sample sizes to deliver efficient learning because one wild swing doesn’t stampede the entire population. In the infinite-memory limit, the randomness washes away in aggregate, and the evolution of market share is governed by a single probability: given today’s shares, what’s the chance a random g-user’s multi-period sample beats a random f-user’s? From that equation, you can read off whether the process marches to zero, marches to one, or stays in between. Here’s a gem: when p is above one-half and you have infinite memory, even sampling N equals one—the smallest possible—delivers efficiency for any omega. Ironically, when memory is absent and inertia is strong, larger samples can actually be counterproductive. Memory restores order by smoothing the shared shock. If you’re hearing all this and thinking, wait, can they really solve the whole dynamic? They’re careful about that. The full, global long-run path of x sub t for every parameter combination is beyond reach. What they give you is a powerful partial picture. Near the endpoints, the linearized drift decides whether those boundaries attract or repel. Using that and some probabilistic scaffolding, they carve the parameter space into regions where you get lock-in, regions where you converge to the better technology, and regions where the process keeps moving and never collapses. In market-diffusion terms, it’s a map with solid landmarks, not a turn-by-turn GPS. What makes the map compelling is how it ties behavior to tangible knobs. Decrease N—talk to fewer people—and you boost conformity. If, on average, one technology is genuinely better and the environment isn’t too wild, that conformity can be socially efficient. Increase N, and you protect diversity. That sounds nice until it means never settling on the right answer. Turn up the popularity weighting, and you intensify herding. Turn on memory, and you reduce the risk that a bad aggregate draw jerks the system into a rut. Each of these adjustments shows up not as vague tendencies but as shifts in the two thresholds, N star of omega and N prime of omega, and changes in how those thresholds compare to the simple payoff ratios built from p. There’s a broader lesson tucked inside the math. We often assume that more information and more communication push societies toward better decisions. Ellison and Fudenberg show a subtler world. Bounded rationality—sampling a few peers, using a simple rule—doesn’t doom us to fads. It can, in fact, be the mechanism that gets a community to the right basin of attraction. The trick is to land in the efficient band—not so little talk that you lock into the wrong thing, and not so much that common shocks and noise keep you churning forever. If you’re imagining extensions—what happens with more than two options, with network structure instead of random mixing, with strategic sampling—the paper hints at the road but doesn’t walk it. That’s for another day. What it nails is a disciplined answer to a very practical question: how much should we talk to each other when we’re learning from experience? Sometimes, less is more. And sometimes, a short memory—or even a long one—makes all the difference in where we end up. This lecture was created by ennepō. Go to ennepo dot A I to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Here’s the puzzle Glenn Ellison and Drew Fudenberg put on the table: if people mostly learn by chatting with a few peers, when does that chatter make everyone pile onto the same choice, and when does it keep a mix alive? The twist that makes their 1995 paper so memorable is this: sometimes talking less pushes people to conform, and under the right conditions, that very conformity steers the whole population to the better action. Conformity isn’t always a bug; it can be a feature.

They build the story in a stripped-down world with two technologies, which they call f and g. Each period, payoffs bounce around because of two kinds of shocks. There’s a common shock that hits everyone in the same direction, and there are idiosyncratic shocks that differ from person to person.

Think weather versus whim. The common shock has two possible values, and there’s a probability p that the state favors g in any given period. The population's state is just one number, x sub t, the fraction currently using g.

People don’t constantly rethink. A fixed share Q of the population reevaluates each period, and when they do, they look out at N randomly chosen peers, compare the average payoff they saw from f and from g in that sample, and switch to whichever had the higher sample average. It’s a very human rule: you look around, see how things went for a handful of others, and copy the better-looking option.

It’s also not fully Bayesian. Because you only switch if you happen to talk to someone on the other side, popular options get observed more, which builds in a subtle popularity weighting.

Under the hood, that simple sampling rule turns the whole system into a Markov process for x sub t. Given x sub t, N, Q, and the common shock for that period, you can write down the probability that a reevaluating person flips to g. In aggregate, that gives you the expected value of x sub t plus one.

If you let N get large, the random sample shares look normal, and the update can be expressed with the standard normal cumulative distribution function—a smooth map that carries today’s share into tomorrow’s. But you don’t need the equation to follow the punchline. The dial that matters most is N, the number of people each person samples.

Turn it down, and one kind of world emerges. Turn it up, and you get a different one.

Ellison and Fudenberg start by looking at what happens near the edges—when a technology is so rare it’s barely hanging on. They prove a clean local result: you can linearize the dynamics around an endpoint and read off the drift. If the expected growth factor near zero is greater than one, then g can’t die out from that neighborhood; the process is repelled from zero.

If the factor is less than one, there’s positive probability you slide all the way down to zero. The trick that makes this work is to analyze the logarithm of the market share near the boundary, where tiny changes in share correspond to big changes in the log. It sounds technical, and it is, but the intuition is simple—ask whether a vanishingly small minority is expected to grow or shrink on its next step. That local criterion ends up anchoring the global picture.

With that tool in hand, they draw the first of their big dividing lines. For each strength of the common shock—call it omega—there’s a critical sample size N star of omega. Below it, the world tends to conformity: one of the two technologies becomes universal with probability one.

Above it, the process does not settle on an endpoint. Shares keep moving, and diversity persists. That N star of omega curve rises with omega.

Stronger common shocks push the boundary up; to get lock-in when the shared environment is volatile, you have to keep samples small. There’s a striking limiting case here that helps build intuition. If you turn off common shocks entirely and leave only idiosyncratic noise, the dynamics become deterministic and collapse to a simple split: for any N of at least two, the market share converges to one-half.

There’s no conformity when weather-like forces disappear and inertia is positive.

Now tilt the playing field so one option is better on average. Suppose p exceeds one-half, so in a typical period g is more likely to win. This is where their analysis sharpens because the model doesn’t just say “the better option wins.” It carves the parameter space into three zones using a function they call Q star of omega and N that comes out of the linearization.

You compare that function to two simple ratios: p over one minus p and its reciprocal. If the ratio p divided by one minus p sits below Q star of omega and N, the population still locks in, but from some starting points it can lock onto the inferior technology—inefficient herding. If Q star of omega and N is sandwiched between those two ratios, everyone ends up on g—efficient social learning.

If Q star of omega and N falls below one minus p divided by p, the system doesn’t converge to an endpoint at all; it stays diverse. It’s an elegant way to say that the average superiority of g helps, but only if the sampling rule lies in the right band.

That band has structure. There isn’t just one boundary; there are two. The lower one is the N star of omega we just met, which marks the pivot from conformity to something else.

There’s also an upper boundary, which we’ll call N prime of omega. When p is over one-half, the region in between—N star of omega to N prime of omega—is where learning is efficient and the population reliably finds and sticks with g. As p rises, that “Goldilocks” band widens.

The lower boundary drops, and the upper boundary rises, meaning a broader swath of sample sizes deliver the right long-run outcome. There’s a tough-love message built in: outside that band, either herding can be wrong, or talking to too many people can keep the system from settling even when one option is better.

You can attach some concrete numbers to these thresholds to get a feel for scale. They look at n, a way to summarize how strong the common shock is relative to idiosyncratic noise. When the common shock is twice as large as individual noise (n equals 2), the critical sample size that matters in their illustrations is on the order of forty.

When the common shock is about the same size as individual noise (n equals 1), the comparable threshold drops to around five. The absolute figures depend on parameterization, but the lesson travels: in volatile shared environments, “ask fewer people” becomes a surprisingly good rule if you care about everyone eventually picking the better technology. Push N very high, and you can get stuck in motion—diversity that never collapses to a single winner and, in their terms, a process that doesn’t even enter a neighborhood of full convergence.

So far, the conversation rule gave each observed payoff equal weight within the sample. Ellison and Fudenberg then ask, what if social influence leans harder on what’s popular? They add a parameter m that tilts sampling probabilities toward whatever is more prevalent.

Think of it as a knob on the popularity amplifier. The three-regime map survives—conformity, efficient learning, diversity—but the borders slide. As Theorem 3 lays out, ratcheting up m generally makes convergence to an endpoint easier.

In some initial conditions, that means the system can get pulled into inefficient lock-in that wouldn’t happen without the extra popularity weighting. The “best” sample size also shifts. The uniformly optimal rule—which in the baseline is N star of omega, the size that guarantees efficiency across all p when p is over one-half—becomes N star of omega and m and rises with m.

In a numerical illustration the authors discuss, with a moderate common shock and p at 0.6, the sweet spot is just a handful of conversations—around three. That detail is less important than the principle: amplify popularity and you don’t need to sample as widely to herd, for better or worse.

Then comes memory. Up to now, people were reacting only to what they and their peers just experienced. Add social memory—either a short, two-period lookback or, in the limit, an infinite one—and the role of the common shock changes.

With even a bit of memory, fluky aggregate conditions get averaged out. That makes it easier for moderate sample sizes to deliver efficient learning because one wild swing doesn’t stampede the entire population. In the infinite-memory limit, the randomness washes away in aggregate, and the evolution of market share is governed by a single probability: given today’s shares, what’s the chance a random g-user’s multi-period sample beats a random f-user’s?

From that equation, you can read off whether the process marches to zero, marches to one, or stays in between. Here’s a gem: when p is above one-half and you have infinite memory, even sampling N equals one—the smallest possible—delivers efficiency for any omega. Ironically, when memory is absent and inertia is strong, larger samples can actually be counterproductive. Memory restores order by smoothing the shared shock.

If you’re hearing all this and thinking, wait, can they really solve the whole dynamic? They’re careful about that. The full, global long-run path of x sub t for every parameter combination is beyond reach.

What they give you is a powerful partial picture. Near the endpoints, the linearized drift decides whether those boundaries attract or repel. Using that and some probabilistic scaffolding, they carve the parameter space into regions where you get lock-in, regions where you converge to the better technology, and regions where the process keeps moving and never collapses.

In market-diffusion terms, it’s a map with solid landmarks, not a turn-by-turn GPS.

What makes the map compelling is how it ties behavior to tangible knobs. Decrease N—talk to fewer people—and you boost conformity. If, on average, one technology is genuinely better and the environment isn’t too wild, that conformity can be socially efficient.

Increase N, and you protect diversity. That sounds nice until it means never settling on the right answer. Turn up the popularity weighting, and you intensify herding.

Turn on memory, and you reduce the risk that a bad aggregate draw jerks the system into a rut. Each of these adjustments shows up not as vague tendencies but as shifts in the two thresholds, N star of omega and N prime of omega, and changes in how those thresholds compare to the simple payoff ratios built from p.

There’s a broader lesson tucked inside the math. We often assume that more information and more communication push societies toward better decisions. Ellison and Fudenberg show a subtler world.

Bounded rationality—sampling a few peers, using a simple rule—doesn’t doom us to fads. It can, in fact, be the mechanism that gets a community to the right basin of attraction. The trick is to land in the efficient band—not so little talk that you lock into the wrong thing, and not so much that common shocks and noise keep you churning forever.

If you’re imagining extensions—what happens with more than two options, with network structure instead of random mixing, with strategic sampling—the paper hints at the road but doesn’t walk it. That’s for another day. What it nails is a disciplined answer to a very practical question: how much should we talk to each other when we’re learning from experience?

Sometimes, less is more. And sometimes, a short memory—or even a long one—makes all the difference in where we end up.

This lecture was created by ennepō.

Go to ennepo dot A I to Discover, Create and Follow the latest research in your field.

Read when you can. Listen when you want to.

More in Social Sciences