Win-Stay-Lose-Learn Promotes Cooperation in the Spatial Prisoner's Dilemma Game

Yongkui Liu, Xiaojie Chen, Zhang Li, Long Wang, Matjaž PercView original
OverviewBalancedriya_rao voice
If you're doing fine, you don't change. That's not a profound insight — it's just how people behave. You stick with what's working. And yet the standard computer simulations that evolutionary game theorists have been running for decades quietly violate that basic fact. In those models, even satisfied and successful cooperators are constantly at risk of abandoning their strategy over the tiniest payoff difference. Liu, Chen, Li, Wang, and Perc examined that assumption and asked what would happen if you fixed it. The answer transformed the fate of cooperation almost entirely. To understand why that matters, you need the prisoner's dilemma. Two players, two choices: cooperate or defect. If both cooperate, each earns the reward, R. If both defect, each earns the punishment, P. But if one defects while the other cooperates, the defector earns temptation, T, and the cooperator earns the sucker's payoff, S — and the payoffs always order as T greater than R greater than P greater than S. That ordering is the trap. No matter what your partner does, defecting earns you more than cooperating. So rational individuals defect, and a population of rational individuals ends up at mutual defection — which is worse for everyone than mutual cooperation would have been. This is why evolutionary game theorists call it a social dilemma: the individually rational move is collectively destructive. In the spatial version of this game, which Liu and colleagues study, players sit on a square lattice and each one interacts only with its four nearest neighbors. The temptation payoff is written as b, ranging between 1 and 2, and everything else is normalized so that mutual cooperation pays 1 and mutual defection pays 0. The higher b is, the stronger the pull toward defection. Now, in most simulations of this game, every player every round compares their payoff to a randomly chosen neighbor's payoff and may adopt that neighbor's strategy. The probability of switching is higher when the neighbor is doing better, but crucially, it's never zero — there's always a chance you switch, even if the payoff difference is marginal. Liu and colleagues point out the problem with this: it means successful cooperators are continuously, round after round, exposed to the possibility of abandoning cooperation over tiny incidental differences. That's not how people actually behave. A satisfied person tends to stay put. So Liu and colleagues introduced what they call the win-stay-lose-learn rule, built around a personal aspiration level. Here's how it works. Each player has an aspiration threshold — on the square lattice with four neighbors, this is simply four times a population-level parameter, A, where A sits between 0 and b. Every round, a player checks their own payoff against their own aspiration. If the payoff meets or exceeds the aspiration, the player is satisfied — they win, they stay. No imitation attempt, no comparison with neighbors, no chance of switching. If the payoff falls below the aspiration, the player is dissatisfied — they lose, so they learn. They pick a random neighbor and adopt that neighbor's strategy with a probability that favors better performers but includes a small noise term, set at 0.1 in the simulations, to allow for occasional errors. The critical departure from standard practice is that the trigger for considering a change is internal: your own payoff versus your own aspiration, not a mandatory round-by-round comparison with whoever happens to be next to you. What does this produce? The results are striking. The parameter that controls everything is the aspiration level, A, and it creates a sweet spot. Set A too low — say A at 0.0 or 0.2 — and almost everyone is satisfied regardless of what they're doing. So defectors hold on just as stubbornly as cooperators, and cooperation plateaus at around 50 and 47 percent, respectively, largely independent of temptation. Set A too high — around 0.8 or 2.0 — and nearly everyone is perpetually dissatisfied; strategies churn constantly, and the model collapses back into standard imitation, where cooperators survive only when temptation, b, is no higher than about 1.05. But in the middle, cooperation flourishes. At A equal to 0.4, the steady-state fraction of cooperators sits around 70 percent for temptation values up to 1.6. At A equal to 0.6, cooperation reaches virtually 100 percent for temptation values up to 1.2. The transitions between these regimes are sharp — Liu and colleagues identify phase-transition points in A at roughly 0.0, 0.25, 0.5, 0.75, and 1.0, arising directly from the payoff arithmetic on a four-neighbor lattice. These results hold up in two independent ways. The simulations themselves track final cooperation fractions across thousands of rounds on large lattices. The team also applies pair approximation — a mathematical technique that models the frequency of neighboring strategy pairs analytically rather than by simulation, tracking how cooperator-cooperator, cooperator-defector, and defector-defector pairs evolve over time. The pair approximation reproduces the qualitative structure of the results and matches the simulations exactly for very low aspirations, providing a theoretical backbone for what the simulations show. What's driving this mechanically? It's the spatial structure combined with the stay-when-satisfied rule. Cooperators who sit together on the lattice earn mutual rewards from their shared interactions. If those rewards meet their aspiration, they freeze — satisfied cooperators simply do not attempt to change strategy. Cooperative clusters become stable islands. Defectors at the boundaries can initially exploit their cooperating neighbors and earn high payoffs from temptation, b, but the cooperators inside the cluster don't budge. Over time, defectors end up surrounded, their exploitable neighbors held in place, and the cluster can expand rather than erode. Under standard imitation rules, those same cooperators would be continuously at risk of switching away. The win-stay rule locks in their success. This also explains the paper's most striking robustness result: cooperation can survive and ultimately dominate even from very small initial fractions. A single isolated cooperator typically cannot hold on — they're too exposed. But a small cluster can. Two neighboring cooperators are stable provided A is at most 0.25; a block of four cooperators can expand and eventually dominate when A stays at or below about 0.5. Liu and colleagues show that the aspiration-tuned rule can rescue cooperation from almost nothing, producing sharp parameter-dependent expansions supported by both simulation and pair-approximation analysis. The paper situates this finding relative to related work. Earlier win-stay rules — such as those studied by Chen and Wang — also restrict strategy updating but do so differently, either by forcing a switch to the opposite strategy when losing or by still requiring pairwise payoff comparisons. The win-stay-lose-learn rule uses an internal aspiration threshold, so often there is no need to compare payoffs with neighbors at all. That difference matters: the aspiration-based gate is what produces stable cooperation at intermediate values of A, and it is what allows the rule to work across a wide range of initial conditions and temptation levels that would defeat traditional approaches. Step back from the mechanics for a moment and consider what the paper is actually arguing. The behavioral assumption baked into standard evolutionary simulations — that players continuously compare payoffs and may switch even over marginal differences — is not just inaccurate. It actively undermines cooperation in the models by keeping successful cooperators in permanent jeopardy. Replace that assumption with something closer to observed behavior — satisfied players hold their ground — and cooperation doesn't just improve modestly. It transforms. The aspiration parameter, A, becomes a dial that determines how much of the population is frozen in place at any given moment, and the optimal setting is one where most satisfied cooperators stay put while dissatisfied players learn from their neighbors. If that principle operates beyond the model — in human organizations, social networks, or biological systems where agents have something like an aspiration or satisfaction threshold — then the implication is direct. Stability, not perpetual optimization, may be what cooperation actually requires to survive. That's the finding Liu and colleagues leave the reader with: a simple aspiration-based rule, requiring only that winners stay, is enough to change the evolutionary fate of cooperation almost entirely. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

If you're doing fine, you don't change. That's not a profound insight — it's just how people behave. You stick with what's working. And yet the standard computer simulations that evolutionary game theorists have been running for decades quietly violate that basic fact. In those models, even satisfied and successful cooperators are constantly at risk of abandoning their strategy over the tiniest payoff difference. Liu, Chen, Li, Wang, and Perc examined that assumption and asked what would happen if you fixed it. The answer transformed the fate of cooperation almost entirely. To understand why that matters, you need the prisoner's dilemma. Two players, two choices: cooperate or defect. If both cooperate, each earns the reward, R. If both defect, each earns the punishment, P. But if one defects while the other cooperates, the defector earns temptation, T, and the cooperator earns the sucker's payoff, S — and the payoffs always order as T greater than R greater than P greater than S. That ordering is the trap. No matter what your partner does, defecting earns you more than cooperating. So rational individuals defect, and a population of rational individuals ends up at mutual defection — which is worse for everyone than mutual cooperation would have been. This is why evolutionary game theorists call it a social dilemma: the individually rational move is collectively destructive.

In the spatial version of this game, which Liu and colleagues study, players sit on a square lattice and each one interacts only with its four nearest neighbors. The temptation payoff is written as b, ranging between 1 and 2, and everything else is normalized so that mutual cooperation pays 1 and mutual defection pays 0. The higher b is, the stronger the pull toward defection. Now, in most simulations of this game, every player every round compares their payoff to a randomly chosen neighbor's payoff and may adopt that neighbor's strategy. The probability of switching is higher when the neighbor is doing better, but crucially, it's never zero — there's always a chance you switch, even if the payoff difference is marginal. Liu and colleagues point out the problem with this: it means successful cooperators are continuously, round after round, exposed to the possibility of abandoning cooperation over tiny incidental differences. That's not how people actually behave. A satisfied person tends to stay put. So Liu and colleagues introduced what they call the win-stay-lose-learn rule, built around a personal aspiration level. Here's how it works. Each player has an aspiration threshold — on the square lattice with four neighbors, this is simply four times a population-level parameter, A, where A sits between 0 and b.

Every round, a player checks their own payoff against their own aspiration. If the payoff meets or exceeds the aspiration, the player is satisfied — they win, they stay. No imitation attempt, no comparison with neighbors, no chance of switching. If the payoff falls below the aspiration, the player is dissatisfied — they lose, so they learn. They pick a random neighbor and adopt that neighbor's strategy with a probability that favors better performers but includes a small noise term, set at 0.1 in the simulations, to allow for occasional errors. The critical departure from standard practice is that the trigger for considering a change is internal: your own payoff versus your own aspiration, not a mandatory round-by-round comparison with whoever happens to be next to you. What does this produce? The results are striking. The parameter that controls everything is the aspiration level, A, and it creates a sweet spot. Set A too low — say A at 0.0 or 0.2 — and almost everyone is satisfied regardless of what they're doing. So defectors hold on just as stubbornly as cooperators, and cooperation plateaus at around 50 and 47 percent, respectively, largely independent of temptation. Set A too high — around 0.8 or 2.0 — and nearly everyone is perpetually dissatisfied; strategies churn constantly, and the model collapses back into standard imitation, where cooperators survive only when temptation, b, is no higher than about 1.05.

But in the middle, cooperation flourishes. At A equal to 0.4, the steady-state fraction of cooperators sits around 70 percent for temptation values up to 1.6. At A equal to 0.6, cooperation reaches virtually 100 percent for temptation values up to 1.2. The transitions between these regimes are sharp — Liu and colleagues identify phase-transition points in A at roughly 0.0, 0.25, 0.5, 0.75, and 1.0, arising directly from the payoff arithmetic on a four-neighbor lattice. These results hold up in two independent ways. The simulations themselves track final cooperation fractions across thousands of rounds on large lattices. The team also applies pair approximation — a mathematical technique that models the frequency of neighboring strategy pairs analytically rather than by simulation, tracking how cooperator-cooperator, cooperator-defector, and defector-defector pairs evolve over time. The pair approximation reproduces the qualitative structure of the results and matches the simulations exactly for very low aspirations, providing a theoretical backbone for what the simulations show. What's driving this mechanically? It's the spatial structure combined with the stay-when-satisfied rule. Cooperators who sit together on the lattice earn mutual rewards from their shared interactions.

If those rewards meet their aspiration, they freeze — satisfied cooperators simply do not attempt to change strategy. Cooperative clusters become stable islands. Defectors at the boundaries can initially exploit their cooperating neighbors and earn high payoffs from temptation, b, but the cooperators inside the cluster don't budge. Over time, defectors end up surrounded, their exploitable neighbors held in place, and the cluster can expand rather than erode. Under standard imitation rules, those same cooperators would be continuously at risk of switching away. The win-stay rule locks in their success. This also explains the paper's most striking robustness result: cooperation can survive and ultimately dominate even from very small initial fractions. A single isolated cooperator typically cannot hold on — they're too exposed. But a small cluster can. Two neighboring cooperators are stable provided A is at most 0.25; a block of four cooperators can expand and eventually dominate when A stays at or below about 0.5. Liu and colleagues show that the aspiration-tuned rule can rescue cooperation from almost nothing, producing sharp parameter-dependent expansions supported by both simulation and pair-approximation analysis.

The paper situates this finding relative to related work. Earlier win-stay rules — such as those studied by Chen and Wang — also restrict strategy updating but do so differently, either by forcing a switch to the opposite strategy when losing or by still requiring pairwise payoff comparisons. The win-stay-lose-learn rule uses an internal aspiration threshold, so often there is no need to compare payoffs with neighbors at all. That difference matters: the aspiration-based gate is what produces stable cooperation at intermediate values of A, and it is what allows the rule to work across a wide range of initial conditions and temptation levels that would defeat traditional approaches. Step back from the mechanics for a moment and consider what the paper is actually arguing. The behavioral assumption baked into standard evolutionary simulations — that players continuously compare payoffs and may switch even over marginal differences — is not just inaccurate. It actively undermines cooperation in the models by keeping successful cooperators in permanent jeopardy. Replace that assumption with something closer to observed behavior — satisfied players hold their ground — and cooperation doesn't just improve modestly. It transforms. The aspiration parameter, A, becomes a dial that determines how much of the population is frozen in place at any given moment, and the optimal setting is one where most satisfied cooperators stay put while dissatisfied players learn from their neighbors.

If that principle operates beyond the model — in human organizations, social networks, or biological systems where agents have something like an aspiration or satisfaction threshold — then the implication is direct. Stability, not perpetual optimization, may be what cooperation actually requires to survive. That's the finding Liu and colleagues leave the reader with: a simple aspiration-based rule, requiring only that winners stay, is enough to change the evolutionary fate of cooperation almost entirely. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Social Sciences