A Dual-Self Model of Impulse Control

Drew Fudenberg, David K. LevineView original
OverviewBalancedbennett voice
Two selves live inside every decision you make — one patient and one impulsive — and they are genuinely in conflict. This is not just a metaphor. It is a formal economic model with a precise mechanism. If that's true, it should explain why you ate the cake, why you're still putting off the dentist, and why memorizing a seven-digit number made both of those worse. That's the claim Drew Fudenberg and David Levine make in their dual-self model of impulse control. One model, one cost, three major behavioral puzzles. Let’s walk through how it works. Classical economic theory assumes impatience is steady. You discount future rewards at a fixed rate, and that rate doesn't depend on whether you're asking about tomorrow versus today or next year versus a year plus a day. Your preferences should be time-consistent. But they aren't. Faced with a choice between a smaller amount now or a larger amount tomorrow, many people take the smaller amount. Those same people, when asked to choose between a larger amount a year from now versus a year and a day from now, will happily wait the extra day. The preference flips depending on whether "now" is in the picture. The economics field's main response to this was quasi-hyperbolic discounting — a mathematical patch that captures the flip by treating today's decision-maker as disproportionately impatient relative to any future decision-maker. It fits the data. But Fudenberg and Levine point out that it's a curve-fitting fix, not a mechanism. It doesn't tell you why preferences flip; it just encodes the flip in the formula. And it creates analytic headaches: model a sequence of such selves and you quickly get multiple equilibria and complicated game-theoretic structures. There's a second puzzle any theory of self-control needs to handle: Rabin's paradox. Experimental subjects show so much risk aversion over small-stakes gambles that, if you extrapolated their behavior using standard expected utility theory, they would have to reject enormous, wildly favorable bets. The degree of aversion to small losses implies an implausible degree of aversion to large risks. Classical utility theory can't reconcile the two. Fudenberg and Levine want a single model that handles time inconsistency and this risk puzzle together. Here's the architecture. The decision problem is a dynamic game between a patient long-run self and a sequence of completely myopic short-run selves, one per period. The short-run self only cares about right now. Before each short-run self acts, the long-run self can intervene — can exert self-control — but doing so costs utility. This is the core mechanism. Willpower is an explicit, costly instrument. The self-control cost depends on two things: the utility of the option actually chosen and the utility of the best forgone alternative. Cost rises with temptation — the bigger the gap between what you want to do and what you choose to do, the more it costs to hold the line. Under the paper's linear specialization, that cost is simply proportional to the utility gap between the best available action and the one actually taken. This linear form keeps things tractable. Equilibrium in the model — what Fudenberg and Levine call short-run perfect Nash equilibrium — requires that every short-run self optimizes given any history, and that the long-run self anticipates all those optimal responses and plans accordingly. They prove that every such equilibrium maps onto the solution of a single reduced-form optimization problem, and conversely. In most applications that problem has a unique solution. So the model is both behaviorally grounded and mathematically clean. Now watch what this machinery generates. Fudenberg and Levine embed the model in a two-subperiod setup: a bank stage, where the long-run self decides how much wealth to keep in the bank and how much cash to carry, and a nightclub stage, where the myopic short-run self can spend whatever cash is in pocket. When preferences are logarithmic — and more generally for constant relative risk aversion utility — the model produces a unique interior solution: a constant savings rate strictly between zero and one. Higher self-control costs reduce the savings rate; greater patience by the long-run self raises it. In the stark version of the setup, an unexpected small cash gift received at the nightclub gets spent entirely — a one hundred percent marginal propensity to consume out of pocket cash — while bank-account gains are mostly saved. That asymmetry matches observed behavior around liquidity. Procrastination follows the same logic. Applied to a stopping-time problem — the kind O'Donoghue and Rabin used with quasi-hyperbolic models — the dual-self model generates the same qualitative result: self-control costs produce excess delay. But the mechanism is cleaner and the equilibrium is unique. The optimal policy takes the form of a cutoff rule: when the current temptation to wait exceeds a threshold, you wait; when it falls below, you act. Expected delay rises with the cost of self-control. One mechanism, same prediction, less analytic baggage. Rabin's paradox gets resolved through the model's two-context structure. Small gambles that are resolved as pocket cash in a tempting situation get evaluated by the short-run self, relative to pocket cash. Large gambles get evaluated relative to lifetime wealth by the long-run self. In the log-utility calibration, subjects who turn down a fifty-fifty bet of losing one hundred dollars versus gaining one hundred and five dollars need not be wildly averse to large favorable bets — they're just exhibiting short-run sensitivity to small cash outcomes. The paper's computations are concrete: a small favorable gain of one hundred five dollars gets rejected when pocket cash is two thousand one hundred dollars or less, while a large favorable gamble gets accepted once lifetime wealth exceeds about four thousand twenty-six dollars. Two different evaluative systems, two very different levels of apparent risk aversion. The paradox dissolves. Then there's the cake experiment. Baba Shiv and Alexander Fedorikhin ran this in nineteen ninety-nine. Subjects memorized either a two-digit or a seven-digit number, then walked to a table with a choice between chocolate cake and fruit salad. When real desserts were present, subjects memorizing the seven-digit number chose cake sixty-three percent of the time. Those memorizing the two-digit number chose cake forty-one percent of the time. This is a statistically significant difference. When only photographs of the desserts were shown, the difference vanished — forty-five percent versus forty-two percent, not significant. What does this mean for the model? Cognitive load didn't change preferences. It raised the cost of exerting self-control. The long-run self had fewer resources available to suppress the short-run pull toward cake. Fudenberg and Levine formalize this by treating external cognitive load as a state variable that shifts the self-control cost upward. The base linear-cost version of the model can be made to match the dessert data, but it's conceptually unsatisfying. Linearity implies that a change in marginal cost from memorization has the same mathematical effect as a change in how attractive the forgone option was — and that misses the psychological point that self-control draws on a limited, depletable resource. So the authors prefer a nonlinear, convex cost function: the more self-control you exert, the faster the cost rises. This extension has a price. The base model with linear costs satisfies the Gul-Pesendorfer axioms — a rigorous axiomatic framework for self-control preferences. One key property is set-betweenness: adding a tempting option to a menu can't make you strictly better off, but it also can't make you worse off than a menu containing only the temptation. Fudenberg and Levine prove that their model satisfies set-betweenness under the linear cost assumption. But nonlinear costs break independence of irrelevant alternatives. The three-dessert example from Dekel, Lipman, and Rustichini makes this vivid: frozen yogurt might be chosen in the presence of ice cream as a strong temptation, but not when the temptation is weaker, producing menu-dependent choices that violate standard axioms. To capture cognitive load properly, the axioms have to be relaxed. The welfare implication runs through all of this. If self-control is costly and context changes that cost, then environment genuinely changes welfare-relevant outcomes. Fudenberg and Levine highlight policy cases — Wertenbroch's stockpiling argument and the drug-criminalization example — where laws that alter transaction costs change the level of temptation and therefore consumption. Choice architecture and nudges, from this perspective, aren't paternalism. They're literally changing the price of exercising self-control. What Fudenberg and Levine built is a single, tightly parameterized mechanism that reproduces what previously required three separate models: quasi-hyperbolic time inconsistency, Rabin's risk paradox, and the cognitive-load effect on temptation. The model connects cleanly to neuroscience and psychology — the paper cites McClure, Laibson, Loewenstein, and Cohen's magnetic resonance imaging evidence — without borrowing their technical language. It treats choice as the interaction of a patient long-run system and myopic short-run systems, and derives behavior from that structure. The broader implication is hard to dismiss: willpower looks less like a fixed character trait and more like a scarce resource with a cost function. And like any resource, it can be taxed, subsidized, and depleted. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Two selves live inside every decision you make — one patient and one impulsive — and they are genuinely in conflict. This is not just a metaphor. It is a formal economic model with a precise mechanism. If that's true, it should explain why you ate the cake, why you're still putting off the dentist, and why memorizing a seven-digit number made both of those worse. That's the claim Drew Fudenberg and David Levine make in their dual-self model of impulse control. One model, one cost, three major behavioral puzzles. Let’s walk through how it works. Classical economic theory assumes impatience is steady. You discount future rewards at a fixed rate, and that rate doesn't depend on whether you're asking about tomorrow versus today or next year versus a year plus a day. Your preferences should be time-consistent. But they aren't. Faced with a choice between a smaller amount now or a larger amount tomorrow, many people take the smaller amount. Those same people, when asked to choose between a larger amount a year from now versus a year and a day from now, will happily wait the extra day. The preference flips depending on whether "now" is in the picture. The economics field's main response to this was quasi-hyperbolic discounting — a mathematical patch that captures the flip by treating today's decision-maker as disproportionately impatient relative to any future decision-maker. It fits the data. But Fudenberg and Levine point out that it's a curve-fitting fix, not a mechanism.

It doesn't tell you why preferences flip; it just encodes the flip in the formula. And it creates analytic headaches: model a sequence of such selves and you quickly get multiple equilibria and complicated game-theoretic structures. There's a second puzzle any theory of self-control needs to handle: Rabin's paradox. Experimental subjects show so much risk aversion over small-stakes gambles that, if you extrapolated their behavior using standard expected utility theory, they would have to reject enormous, wildly favorable bets. The degree of aversion to small losses implies an implausible degree of aversion to large risks. Classical utility theory can't reconcile the two. Fudenberg and Levine want a single model that handles time inconsistency and this risk puzzle together. Here's the architecture. The decision problem is a dynamic game between a patient long-run self and a sequence of completely myopic short-run selves, one per period. The short-run self only cares about right now. Before each short-run self acts, the long-run self can intervene — can exert self-control — but doing so costs utility. This is the core mechanism. Willpower is an explicit, costly instrument.

The self-control cost depends on two things: the utility of the option actually chosen and the utility of the best forgone alternative. Cost rises with temptation — the bigger the gap between what you want to do and what you choose to do, the more it costs to hold the line. Under the paper's linear specialization, that cost is simply proportional to the utility gap between the best available action and the one actually taken. This linear form keeps things tractable. Equilibrium in the model — what Fudenberg and Levine call short-run perfect Nash equilibrium — requires that every short-run self optimizes given any history, and that the long-run self anticipates all those optimal responses and plans accordingly. They prove that every such equilibrium maps onto the solution of a single reduced-form optimization problem, and conversely. In most applications that problem has a unique solution. So the model is both behaviorally grounded and mathematically clean. Now watch what this machinery generates. Fudenberg and Levine embed the model in a two-subperiod setup: a bank stage, where the long-run self decides how much wealth to keep in the bank and how much cash to carry, and a nightclub stage, where the myopic short-run self can spend whatever cash is in pocket. When preferences are logarithmic — and more generally for constant relative risk aversion utility — the model produces a unique interior solution: a constant savings rate strictly between zero and one.

Higher self-control costs reduce the savings rate; greater patience by the long-run self raises it. In the stark version of the setup, an unexpected small cash gift received at the nightclub gets spent entirely — a one hundred percent marginal propensity to consume out of pocket cash — while bank-account gains are mostly saved. That asymmetry matches observed behavior around liquidity. Procrastination follows the same logic. Applied to a stopping-time problem — the kind O'Donoghue and Rabin used with quasi-hyperbolic models — the dual-self model generates the same qualitative result: self-control costs produce excess delay. But the mechanism is cleaner and the equilibrium is unique. The optimal policy takes the form of a cutoff rule: when the current temptation to wait exceeds a threshold, you wait; when it falls below, you act. Expected delay rises with the cost of self-control. One mechanism, same prediction, less analytic baggage. Rabin's paradox gets resolved through the model's two-context structure. Small gambles that are resolved as pocket cash in a tempting situation get evaluated by the short-run self, relative to pocket cash. Large gambles get evaluated relative to lifetime wealth by the long-run self.

In the log-utility calibration, subjects who turn down a fifty-fifty bet of losing one hundred dollars versus gaining one hundred and five dollars need not be wildly averse to large favorable bets — they're just exhibiting short-run sensitivity to small cash outcomes. The paper's computations are concrete: a small favorable gain of one hundred five dollars gets rejected when pocket cash is two thousand one hundred dollars or less, while a large favorable gamble gets accepted once lifetime wealth exceeds about four thousand twenty-six dollars. Two different evaluative systems, two very different levels of apparent risk aversion. The paradox dissolves. Then there's the cake experiment. Baba Shiv and Alexander Fedorikhin ran this in nineteen ninety-nine. Subjects memorized either a two-digit or a seven-digit number, then walked to a table with a choice between chocolate cake and fruit salad. When real desserts were present, subjects memorizing the seven-digit number chose cake sixty-three percent of the time. Those memorizing the two-digit number chose cake forty-one percent of the time. This is a statistically significant difference. When only photographs of the desserts were shown, the difference vanished — forty-five percent versus forty-two percent, not significant. What does this mean for the model? Cognitive load didn't change preferences. It raised the cost of exerting self-control.

The long-run self had fewer resources available to suppress the short-run pull toward cake. Fudenberg and Levine formalize this by treating external cognitive load as a state variable that shifts the self-control cost upward. The base linear-cost version of the model can be made to match the dessert data, but it's conceptually unsatisfying. Linearity implies that a change in marginal cost from memorization has the same mathematical effect as a change in how attractive the forgone option was — and that misses the psychological point that self-control draws on a limited, depletable resource. So the authors prefer a nonlinear, convex cost function: the more self-control you exert, the faster the cost rises. This extension has a price. The base model with linear costs satisfies the Gul-Pesendorfer axioms — a rigorous axiomatic framework for self-control preferences. One key property is set-betweenness: adding a tempting option to a menu can't make you strictly better off, but it also can't make you worse off than a menu containing only the temptation.

Fudenberg and Levine prove that their model satisfies set-betweenness under the linear cost assumption. But nonlinear costs break independence of irrelevant alternatives. The three-dessert example from Dekel, Lipman, and Rustichini makes this vivid: frozen yogurt might be chosen in the presence of ice cream as a strong temptation, but not when the temptation is weaker, producing menu-dependent choices that violate standard axioms. To capture cognitive load properly, the axioms have to be relaxed. The welfare implication runs through all of this. If self-control is costly and context changes that cost, then environment genuinely changes welfare-relevant outcomes. Fudenberg and Levine highlight policy cases — Wertenbroch's stockpiling argument and the drug-criminalization example — where laws that alter transaction costs change the level of temptation and therefore consumption. Choice architecture and nudges, from this perspective, aren't paternalism. They're literally changing the price of exercising self-control.

What Fudenberg and Levine built is a single, tightly parameterized mechanism that reproduces what previously required three separate models: quasi-hyperbolic time inconsistency, Rabin's risk paradox, and the cognitive-load effect on temptation. The model connects cleanly to neuroscience and psychology — the paper cites McClure, Laibson, Loewenstein, and Cohen's magnetic resonance imaging evidence — without borrowing their technical language. It treats choice as the interaction of a patient long-run system and myopic short-run systems, and derives behavior from that structure. The broader implication is hard to dismiss: willpower looks less like a fixed character trait and more like a scarce resource with a cost function. And like any resource, it can be taxed, subsidized, and depleted. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Decision Sciences