The anatomy of choiceactive inference and agency

Karl Friston, Philipp Schwartenbeck, Thomas H. B. FitzGerald, Michael Moutoussis, Timothy E.J. Behrens, Raymond J. DolanView original
OverviewBalancedalloy voice
Let's flip the usual picture of decision-making. Not "What action maximizes reward?" but "What action makes my beliefs about the future and my goals line up?" That's the move Karl Friston and colleagues make. In their telling, agency is an inference problem. You carry a model of how the world unfolds and a set of beliefs about which states you should end up in, and you act to minimize the gap between those two distributions. Technically, that gap is a Kullback-Leibler, or K L, divergence. Conceptually, it's the mismatch between the futures you can actually reach and the futures you want to make likely. Action becomes the downstream consequence of inference. This idea lives inside a larger principle: keep surprise low. If you minimize surprise over time—surprise in the Shannon sense, the negative log probability of your observations—you will, under mild assumptions about how often states recur, spend your life in a small set of familiar, self-organized states. Sounds biological. Exact Bayesian inference about the hidden causes of sensations is intractable, though, so the trick is to optimize a proxy called variational free energy. Free energy is an upper bound on surprise; push it down and you both reduce expected surprise and increase model evidence. Three faces of the same coin show up here. You can write free energy as expected Gibbs energy minus posterior entropy. You can see it as a nonnegative bound on surprise. Or you can see it as the objective that, when minimized, maximizes expected log likelihood while pulling your approximate posterior toward your prior in a controlled way. The punchline is that perception and action are two sides of the same inference: you update beliefs to explain what you see, and you act to make what you'll see easier to explain. To make this more than a slogan, they build a generative model. Hidden states of the world generate observations; a subset of those hidden states are control states—basically, your own actions—that modulate how world states transition. Observations given states are parameterized by a matrix often called A. State transitions under each action are captured by B of u, one matrix per control state. Priors over where you start and where you ought to end up are encoded by vectors c and d. And here's the crucial piece: prior beliefs over whole policies—sequences of actions you might take—follow a Boltzmann form. In plainer language, the more a policy pushes the distribution of final states to look like your goal distribution, the more likely you believe that policy is before seeing anything new. There's a precision term tucked in there too, a positive scalar called gamma that plays the role of an inverse temperature. But it's not an external knob. It's a hidden variable to be inferred, with a gamma prior—shape eight, rate one in their simulations. Confidence is part of the inference. Now, if you're thinking, "Isn't this just expected utility in disguise?"—sometimes, yes. But the route matters. In this setup, what looks like utility falls out of the inference itself. The probability of a policy from your current state obeys a Boltzmann-like rule: take the value of that policy and weight it by precision. In symbols you don't have to see to feel: the log probability of a policy equals gamma times its value. And the value decomposes cleanly. One term is intrinsic: the entropy over final states given that policy. That's an exploration bonus. The other is extrinsic: the expected log probability of those final states under your prior goals—what you'd normally call utility. So a policy is good if it leaves wiggle room to learn or if it drives you toward prized outcomes. Often both. When your beliefs about the future are sharp, that extrinsic term dominates and you look like a classic expected-utility maximizer. When they're flat, the entropy term takes the lead, and you seek information the way a maximum-entropy agent would. Exploration versus exploitation here isn't a dial you set; it's a balance your beliefs determine. There are a few more moving parts worth grounding. Policies are enumerated—say there are K of them—and for each current state, you can store their values in a K by J matrix, Q. That symbol Q here is not the Q of reinforcement learning; it's just a table of policy values from different current states. There's also a mapping, T of pi, that tells you, policy by policy, the probability distribution over final states you'll reach from a given starting state. Under the prior, the best policy is precisely the one that makes T of pi's final-state distribution closest to your goal distribution. That's the K L story you met at the top, restated in the nuts and bolts of the model. Inference proceeds by inverting this generative model with variational Bayes. To keep the math tractable, they use a mean-field factorization: split the hidden variables into blocks—past hidden states, future control states, and precision—and let each block update based on the others' current expectations. Time gets two speeds. On a fast timescale, you update beliefs about what's going on right now given the latest observations. On a slower timescale, you update beliefs about which policies are plausible and how precise your beliefs should be as actions and outcomes accumulate. The posterior over precision has a convenient conjugate form, so you keep that gamma prior's shape and dial its rate up or down with experience. It's a Bayesian thermostat. There's a subtlety with policies. Outcomes depend on whole sequences of actions, so you can't just factor the marginal over future control states into independent time steps without breaking things. The workaround is pragmatic: restrict attention to plausible policy sequences—no impossible backtracks in an irreversible process, for example—and keep the joint over those candidates. This containment keeps the posterior expressive enough to care about history but small enough to compute with. How do the actual updates look? Softmax all the way down. The expectation of a hidden state becomes a softmax over a mixture of signals: what the A matrix says about observations and what the B matrices say about transitions. The expectation over policies becomes a softmax over their values scaled by the current precision. And the precision itself updates from the inside out, using the current expectations about states and choices. You see a predicted monotonic relationship: as the attainable value goes up, the expected precision rises. When there's little value to be had, precision relaxes toward its prior level—remember that shape parameter of eight—reflecting broader, more exploratory beliefs. Precision here plays a dual role. It sharpens action selection—like an inverse temperature in a softmax choice rule—and it biases perceptual inference toward futures that make sense under policies you take seriously. Confidence about what you will do and about what you will encounter are entangled. This has some immediate behavioral consequences. Because that intrinsic term in policy value is literally the entropy over final states, you expect agents to prefer policies that keep options open when extrinsic utilities are tied. That's a concrete, testable prediction: if two outcomes are equally valued, choices should tilt toward higher-entropy futures. You also get softmax-like, "quantal response" behavior without postulating noise; stochasticity in choice comes from the posterior over policies and the current precision. And you get what psychologists would call optimism bias as a side effect of selecting policies that make desired futures more probable—your beliefs about what will happen and your beliefs about what you will do reinforce each other through the same inferential loop. All of this sits neatly next to related literatures. In control theory, K L control introduces a regularization term that penalizes deviations from a baseline dynamics; here, the K L lives between the distribution of attainable futures under a policy and the distribution you aim for. In predictive coding, perception is the suppression of prediction errors by updating beliefs; active inference glues action onto that same architecture, shrinking errors not just by belief updates but by sampling the world in ways your model expects. And because the precision term behaves like an inverse temperature that ramps up when valuable outcomes are within reach, it's tempting to see neuromodulatory systems—dopamine, in particular—as carrying a precision-like signal. Friston and colleagues sketch that mapping in broad strokes and point to convergences with known dopaminergic dynamics. But they're careful: it's a suggestive narrative, not settled anatomy. Let me pause and summarize the equation that anchors the behavioral story, but in words. The log probability of choosing a particular sequence of actions from your current state equals a precision parameter times the value of that sequence. And the value equals two pieces added together: first, the entropy of the distribution over final states if you took that sequence—how many futures it keeps alive—and second, the expected log probability of those final states under your prior preferences—the extrinsic utility. That one relationship explains why uncertainty can feel rewarding, why confidence tightens choices, and why "rational" behavior emerges as a special case when you already know how the world will respond. The computational scheme that implements this is disciplined but approximate. Exact Bayesian inversion is out of reach, so the model assumes a factorized posterior, works in discrete time, and often simplifies the policy space. The updates are derived from variational energies—think of them as negative Gibbs energies—computed for each block of hidden variables by taking expectations over the rest. Iterate these updates until things settle, and you've got Bayesian estimates of the hidden state of the world, the policies worth pursuing, and how confident to be. It's a coherent story about bounded rationality: agents minimize a free-energy bound that trades accuracy against complexity, and the "temperature" of their choices is not arbitrary noise but an inferred quantity tied to the attainable value in their environment. What about timing and value over time—discounting, the thrill of a big offer, the letdown after a miss? In simulations, as they describe, precision tends to climb after evidence of high value appears and ease off when the value landscape flattens. That dynamic can look like temporal discounting emerging from the structure of the generative model and the way precision is updated, not from a hand-coded discount factor. It's an appealing bridge from sequential decision phenomena to the math we've been walking through. There are acknowledged limitations. The mean-field assumptions break dependencies you might care about. The discrete-time setup is a convenience, not a law of nature. And while it's enticing to read the precision updates as dopamine and the message passing as cortical-subcortical loops, the authors flag the neurobiological mapping as provisional. The safer ground is computational: if you accept the free-energy framework, values, curiosity, softmax choice rules, and bounded rationality all fall out of the same inference engine. Two final notes before we close. First, the framework makes concrete predictions you can take to the lab: entropy-seeking when utilities tie, softmax response curves shaped by inferred precision rather than fixed noise, and a mutual dependence of beliefs about states and about one's own future actions. Second, there are areas we haven't covered here—like a detailed mapping between mean-field updates and specific brain circuits, or how precision signals might drive accept-or-wait choices in neuroeconomic tasks—because the material at hand didn't include that evidence. When those links are pinned down with data, the story will either tighten or bend. For now, the payoff is this reframing: agency as inference. You don't act because you've computed a separate utility and then bolted on exploration. You act because you're constantly aligning your anticipated futures with your preferred ones, under uncertainty, with confidence that itself is estimated. And when you do that, the familiar patterns of behavior—curiosity, decisiveness, even bias—stop being add-ons. They're what principled inference looks like when it has a body.

Let's flip the usual picture of decision-making. Not "What action maximizes reward?" but "What action makes my beliefs about the future and my goals line up?" That's the move Karl Friston and colleagues make. In their telling, agency is an inference problem.

You carry a model of how the world unfolds and a set of beliefs about which states you should end up in, and you act to minimize the gap between those two distributions. Technically, that gap is a Kullback-Leibler, or K L, divergence. Conceptually, it's the mismatch between the futures you can actually reach and the futures you want to make likely. Action becomes the downstream consequence of inference.

This idea lives inside a larger principle: keep surprise low. If you minimize surprise over time—surprise in the Shannon sense, the negative log probability of your observations—you will, under mild assumptions about how often states recur, spend your life in a small set of familiar, self-organized states. Sounds biological.

Exact Bayesian inference about the hidden causes of sensations is intractable, though, so the trick is to optimize a proxy called variational free energy. Free energy is an upper bound on surprise; push it down and you both reduce expected surprise and increase model evidence. Three faces of the same coin show up here.

You can write free energy as expected Gibbs energy minus posterior entropy. You can see it as a nonnegative bound on surprise. Or you can see it as the objective that, when minimized, maximizes expected log likelihood while pulling your approximate posterior toward your prior in a controlled way.

The punchline is that perception and action are two sides of the same inference: you update beliefs to explain what you see, and you act to make what you'll see easier to explain.

To make this more than a slogan, they build a generative model. Hidden states of the world generate observations; a subset of those hidden states are control states—basically, your own actions—that modulate how world states transition. Observations given states are parameterized by a matrix often called A.

State transitions under each action are captured by B of u, one matrix per control state. Priors over where you start and where you ought to end up are encoded by vectors c and d. And here's the crucial piece: prior beliefs over whole policies—sequences of actions you might take—follow a Boltzmann form.

In plainer language, the more a policy pushes the distribution of final states to look like your goal distribution, the more likely you believe that policy is before seeing anything new. There's a precision term tucked in there too, a positive scalar called gamma that plays the role of an inverse temperature. But it's not an external knob.

It's a hidden variable to be inferred, with a gamma prior—shape eight, rate one in their simulations. Confidence is part of the inference.

Now, if you're thinking, "Isn't this just expected utility in disguise?"—sometimes, yes. But the route matters. In this setup, what looks like utility falls out of the inference itself.

The probability of a policy from your current state obeys a Boltzmann-like rule: take the value of that policy and weight it by precision. In symbols you don't have to see to feel: the log probability of a policy equals gamma times its value. And the value decomposes cleanly.

One term is intrinsic: the entropy over final states given that policy. That's an exploration bonus. The other is extrinsic: the expected log probability of those final states under your prior goals—what you'd normally call utility.

So a policy is good if it leaves wiggle room to learn or if it drives you toward prized outcomes. Often both. When your beliefs about the future are sharp, that extrinsic term dominates and you look like a classic expected-utility maximizer.

When they're flat, the entropy term takes the lead, and you seek information the way a maximum-entropy agent would. Exploration versus exploitation here isn't a dial you set; it's a balance your beliefs determine.

There are a few more moving parts worth grounding. Policies are enumerated—say there are K of them—and for each current state, you can store their values in a K by J matrix, Q. That symbol Q here is not the Q of reinforcement learning; it's just a table of policy values from different current states.

There's also a mapping, T of pi, that tells you, policy by policy, the probability distribution over final states you'll reach from a given starting state. Under the prior, the best policy is precisely the one that makes T of pi's final-state distribution closest to your goal distribution. That's the K L story you met at the top, restated in the nuts and bolts of the model.

Inference proceeds by inverting this generative model with variational Bayes. To keep the math tractable, they use a mean-field factorization: split the hidden variables into blocks—past hidden states, future control states, and precision—and let each block update based on the others' current expectations. Time gets two speeds.

On a fast timescale, you update beliefs about what's going on right now given the latest observations. On a slower timescale, you update beliefs about which policies are plausible and how precise your beliefs should be as actions and outcomes accumulate. The posterior over precision has a convenient conjugate form, so you keep that gamma prior's shape and dial its rate up or down with experience. It's a Bayesian thermostat.

There's a subtlety with policies. Outcomes depend on whole sequences of actions, so you can't just factor the marginal over future control states into independent time steps without breaking things. The workaround is pragmatic: restrict attention to plausible policy sequences—no impossible backtracks in an irreversible process, for example—and keep the joint over those candidates.

This containment keeps the posterior expressive enough to care about history but small enough to compute with.

How do the actual updates look? Softmax all the way down. The expectation of a hidden state becomes a softmax over a mixture of signals: what the A matrix says about observations and what the B matrices say about transitions.

The expectation over policies becomes a softmax over their values scaled by the current precision. And the precision itself updates from the inside out, using the current expectations about states and choices. You see a predicted monotonic relationship: as the attainable value goes up, the expected precision rises.

When there's little value to be had, precision relaxes toward its prior level—remember that shape parameter of eight—reflecting broader, more exploratory beliefs. Precision here plays a dual role. It sharpens action selection—like an inverse temperature in a softmax choice rule—and it biases perceptual inference toward futures that make sense under policies you take seriously.

Confidence about what you will do and about what you will encounter are entangled.

This has some immediate behavioral consequences. Because that intrinsic term in policy value is literally the entropy over final states, you expect agents to prefer policies that keep options open when extrinsic utilities are tied. That's a concrete, testable prediction: if two outcomes are equally valued, choices should tilt toward higher-entropy futures.

You also get softmax-like, "quantal response" behavior without postulating noise; stochasticity in choice comes from the posterior over policies and the current precision. And you get what psychologists would call optimism bias as a side effect of selecting policies that make desired futures more probable—your beliefs about what will happen and your beliefs about what you will do reinforce each other through the same inferential loop.

All of this sits neatly next to related literatures. In control theory, K L control introduces a regularization term that penalizes deviations from a baseline dynamics; here, the K L lives between the distribution of attainable futures under a policy and the distribution you aim for. In predictive coding, perception is the suppression of prediction errors by updating beliefs; active inference glues action onto that same architecture, shrinking errors not just by belief updates but by sampling the world in ways your model expects.

And because the precision term behaves like an inverse temperature that ramps up when valuable outcomes are within reach, it's tempting to see neuromodulatory systems—dopamine, in particular—as carrying a precision-like signal. Friston and colleagues sketch that mapping in broad strokes and point to convergences with known dopaminergic dynamics. But they're careful: it's a suggestive narrative, not settled anatomy.

Let me pause and summarize the equation that anchors the behavioral story, but in words. The log probability of choosing a particular sequence of actions from your current state equals a precision parameter times the value of that sequence. And the value equals two pieces added together: first, the entropy of the distribution over final states if you took that sequence—how many futures it keeps alive—and second, the expected log probability of those final states under your prior preferences—the extrinsic utility.

That one relationship explains why uncertainty can feel rewarding, why confidence tightens choices, and why "rational" behavior emerges as a special case when you already know how the world will respond.

The computational scheme that implements this is disciplined but approximate. Exact Bayesian inversion is out of reach, so the model assumes a factorized posterior, works in discrete time, and often simplifies the policy space. The updates are derived from variational energies—think of them as negative Gibbs energies—computed for each block of hidden variables by taking expectations over the rest.

Iterate these updates until things settle, and you've got Bayesian estimates of the hidden state of the world, the policies worth pursuing, and how confident to be. It's a coherent story about bounded rationality: agents minimize a free-energy bound that trades accuracy against complexity, and the "temperature" of their choices is not arbitrary noise but an inferred quantity tied to the attainable value in their environment.

What about timing and value over time—discounting, the thrill of a big offer, the letdown after a miss? In simulations, as they describe, precision tends to climb after evidence of high value appears and ease off when the value landscape flattens. That dynamic can look like temporal discounting emerging from the structure of the generative model and the way precision is updated, not from a hand-coded discount factor.

It's an appealing bridge from sequential decision phenomena to the math we've been walking through.

There are acknowledged limitations. The mean-field assumptions break dependencies you might care about. The discrete-time setup is a convenience, not a law of nature.

And while it's enticing to read the precision updates as dopamine and the message passing as cortical-subcortical loops, the authors flag the neurobiological mapping as provisional. The safer ground is computational: if you accept the free-energy framework, values, curiosity, softmax choice rules, and bounded rationality all fall out of the same inference engine.

Two final notes before we close. First, the framework makes concrete predictions you can take to the lab: entropy-seeking when utilities tie, softmax response curves shaped by inferred precision rather than fixed noise, and a mutual dependence of beliefs about states and about one's own future actions. Second, there are areas we haven't covered here—like a detailed mapping between mean-field updates and specific brain circuits, or how precision signals might drive accept-or-wait choices in neuroeconomic tasks—because the material at hand didn't include that evidence.

When those links are pinned down with data, the story will either tighten or bend.

For now, the payoff is this reframing: agency as inference. You don't act because you've computed a separate utility and then bolted on exploration. You act because you're constantly aligning your anticipated futures with your preferred ones, under uncertainty, with confidence that itself is estimated.

And when you do that, the familiar patterns of behavior—curiosity, decisiveness, even bias—stop being add-ons. They're what principled inference looks like when it has a body.

More in Arts and Humanities