Attention, Uncertainty, and Free-Energy

Harriet Feldman, Karl FristonView original
OverviewBalancedalloy voice
Picture attention not as a spotlight you shine, but as a bet you place. When you think a channel is trustworthy, you raise its volume. When you doubt it, you turn it down. That's the core move in the free-energy view of attention from Karl Friston and colleagues: attention is inference about precision, which is the brain's estimate of how reliable a sensory prediction error is. Precision here is just the flip side of variance. High precision means low noise. And in neural hardware, that precision is carried by the gain on the neurons that report mismatch—prediction error units—so attention becomes a story about tuning their postsynaptic gain. Why tune it? Because the brain is constantly trying to explain away its sensory data with a generative model, and it does that by minimizing a quantity called variational free energy. You can think of free energy as an upper bound on surprise, or as a running log-evidence for your current model of the world. Drive it down, and your predictions get better. Under simplifying assumptions, that minimization looks like turning down prediction error. Here's the trick: it's not just the states of the world that get optimized—like where the object is and what it is—it's also the precision with which you think those states are encoded. If precision is itself state-dependent, then attention comes along for free. The brain starts allocating gain to the channels it believes are most reliable, and all of a sudden you see hallmark attention effects—faster detection at cued locations and suppression of distractors—those early and late blips in event-related potentials emerge from the same inferential machinery. Under the hood sits a hierarchical generative model. Causes live at higher levels, states live below them, and the two together generate the sensory stream through a set of nonlinear mappings. The brain keeps a running belief about these quantities—statisticians would call it a recognition density—and a standard move is to approximate that belief as Gaussian, with a mean and a covariance. Flip the covariance and you get precision, and that precision shows up everywhere: it scales how strongly prediction errors pull on your beliefs. There's a compact statement of this idea in the math. The free energy, when you assume a Gaussian belief, is dominated by a simple quadratic in the prediction errors for data, for how fast states are changing, and for parameters; the curvature of that quadratic sets the precision. Match your precision to that curvature and, conveniently, you can update the mean as if it were the only thing that mattered. Now, not all noise is created equal. In these models, the amplitude of random fluctuations—the sensory noise—depends on context. The log-precision of that noise is a function of hidden states and higher-level causes. In plainer words, what you think is going on in the scene controls how noisy you expect each channel to be. Those expectations modulate the gain of the forward-going error units, the ones often linked to superficial pyramidal cells that show up in electroencephalography. Slower, neuromodulatory parameters tune this precision too; think of cholinergic systems adjusting the overall gain landscape. Learning those parameters uses the same gradient rules as learning the content of the generative model, so attention and learning aren't separate hacks—they're two faces of the same updating scheme. The day-to-day dynamics look like message passing. Bottom-up and lateral error signals tug beliefs toward better explanations, while top-down predictions constantly reshape what counts as a good prior. The algorithmic gloss is a kind of generalized filtering: linearize locally, use the Jacobian to steer your updates, and let curvature terms supply the precision that scales each error. It's mathy, but the behavior is intuitive. If a channel is precise, its errors count for more and your beliefs move quickly in response. If it's noisy, you ignore it. That's attention, rendered as Bayes. Bring this down to a classic task: Posner cueing. You get a cue that's valid most of the time—about four out of five in the foundational studies—and then a target flashes left or right. Valid cues speed you up; invalid ones slow you down. In the precision-weighting story, the cue is nothing mystical. It just sets up a high-precision context for the cued location. When the target arrives, the prediction errors from that channel are amplified and push beliefs fast. Errors from the uncued side are down-weighted and drift. The precision context itself has inertia, so some of the bias persists after the cue, then decays. To show this in action, Friston's group built a minimal two-level model. The first level encodes visual inputs. The second level encodes the causes: one for a left target, one for a right target, and one for the central cue. Hidden states at the first level implement a complementary precision mechanism, like a seesaw—if the left gets high precision, the right gives some up. The sensory noise is modeled with a log-precision that depends on these states. If the signal is strong, the model expects less noise, consistent with Weber's law. There's also a neat link to Poisson-like sensory variability: represent the signal as the log of a rate, and the precision becomes proportional to that rate. The cue itself is a brief Gaussian bump in time, on the order of a few dozen milliseconds, while the target is a slightly longer biphasic pulse. Hidden states accumulate and decay over hundreds of milliseconds, so the attentional set lingers. What falls out matches the psychology. In valid trials, the target's representation rides the high-precision channel, and the posterior confidence that a target is present rises quickly; in invalid trials, it ambles. If you take a decision threshold—say you'll respond when posterior confidence hits a set point—you get a reaction time advantage for valid cues of about twenty milliseconds in the model, which is right in the empirical ballpark. The time course tracks intuition too. Cue-driven effects bloom after a couple of hundred milliseconds and then relax as the precision context fades. Neutral cues, which crank up alertness without biasing space, yield a small timing benefit without favoring one side. The electrophysiology lines up as well. Simulated event-related potentials show two distinct epochs. Early, valid cues yield bigger, sharper sensory responses—what people label P1 or N1 in the waveform—consistent with a gain boost on sensory prediction errors. Later, invalid trials kick up larger errors at the level of context, which looks like the brain updating its model of where precision should be after being surprised—think of the P3. Mangun and Hillyard saw this split in real data decades ago, and the match to that pattern is a strong check that the mechanism is pointing the right way. The model also produces slow, anticipatory baseline shifts—contingent negative variation—that reflect the inferred precision in the attended channel before the target even appears. Put two targets on the screen at once and ask the system to pick. In the classic biased competition view from Desimone and Duncan, stimuli that share a receptive field fight it out, with mutual suppression that grows as you move up the visual hierarchy and receptive fields widen. Imaging in humans shows the same logic in populations: further along the ventral stream, the farther apart two items can be and still suppress each other. Top-down goal signals and bottom-up salience both tilt the contest. It's a tidy, well-supported story. What the precision view does is, it doesn't throw that story out; it explains it. Only one precision context can be in force at a time. If the system infers that left is precise now, then left-channel errors get cranked up in their influence and right-channel errors lose steam. When you present two targets simultaneously, the attended one elicits a strong, narrow representation of its cause. The unattended one is scrubbed down, often to roughly half the response of the attended item. Add a valid cue and the unattended channel's expectation drops further—on the order of twenty percent—with wider uncertainty, because the precision context has shifted even more decisively. No explicit lateral inhibition needs to be wired into the generative model; the competition comes for free from inferred precision and the nonlinearity that precision injects into the message passing. Behaviorally, that precision reallocation shows up as the standard speed and accuracy trade-offs. Valid cues buy you speed without much accuracy loss up to a point; invalid cues make you cautious or late, depending on where you set your threshold. In electrophysiology, you get the early sensory boost when your model's precision matches the target, and the late context update when it doesn't. Luck and colleagues saw the early sensory enhancement track attention in human recordings. The later, larger responses to invalid events look like a context surprise—the brain realizing its precision bet was off and recalibrating. There's neurobiology ready to plug into these dials. Gamma band synchronization—the fast rhythms between thirty and one hundred hertz—has long been tied to effective gain on synaptic inputs. Cholinergic systems modulate postsynaptic gain and adaptation and are prime candidates for implementing precision control. Put those facts next to the model's claim that attention is postsynaptic gain on prediction error units, and you have a plausible mechanism bridging computation and tissue. Baseline shifts, early and late event-related components, gamma rhythms, cholinergic tone—they all fall into place as ways the brain encodes, exploits, and updates precision. If you like equations stated in words, here's the essential one the simulations lean on. Sensory noise is Gaussian with a variance that is the exponential of the negative log-precision. The log-precision itself depends on hidden states and a parameter that sets how strongly context moves precision around. Because the model represents the sensory signal as a log rate, the precision of that signal equals the underlying rate. That simple relationship makes precision an explicit, movable resource you can reallocate on the fly as the context changes. There are honest limitations. The model is minimal: two locations, two levels, no explicit lateral wiring, and no attempt to spell out the spatial textures of receptive fields across the cortex. All the attentional boosting appears only when you invert the model—when you run the inference—not because you hard coded gain into the sensory inputs. That's a feature philosophically, but it leaves open questions about how best to map the abstract nodes onto real circuits. The slower learning of precision over longer timescales is sketched but not yet fit to data, and there's a lot of room for formal comparison to competing models on real datasets. Even with those caveats, the through-line is compelling. Start with a brain trying to minimize surprise by explaining its inputs. Let that brain learn not just what's out there but how reliable each channel is in the current context. Then watch attention fall out of the math: valid cue benefits of around twenty milliseconds, early sensory amplification when your precision bet is right, late context updates when it's wrong, and mutual suppression between neighbors when you cram the scene. Desimone and Duncan's competition isn't replaced; it's recast as the inevitable consequence of reallocating precision in a hierarchical inference engine. Where does this go next? The predictions are specific enough to test. Tie reaction time curves and event-related potentials to model fits, not just qualitative matches. Manipulate gamma rhythms or cholinergic tone and watch the inferred precision parameters move. Fit the precision learning knobs on timescales of a few hundred trials and see if the model tracks human adjustment to changing cue validities. If attention is the brain's bet on which errors to trust, then every cue, every distractor, and every burst of acetylcholine is a nudge on the odds. And the free energy view gives us the book the brain is keeping as it lays those bets.

Picture attention not as a spotlight you shine, but as a bet you place. When you think a channel is trustworthy, you raise its volume. When you doubt it, you turn it down.

That's the core move in the free-energy view of attention from Karl Friston and colleagues: attention is inference about precision, which is the brain's estimate of how reliable a sensory prediction error is. Precision here is just the flip side of variance. High precision means low noise.

And in neural hardware, that precision is carried by the gain on the neurons that report mismatch—prediction error units—so attention becomes a story about tuning their postsynaptic gain.

Why tune it? Because the brain is constantly trying to explain away its sensory data with a generative model, and it does that by minimizing a quantity called variational free energy. You can think of free energy as an upper bound on surprise, or as a running log-evidence for your current model of the world.

Drive it down, and your predictions get better. Under simplifying assumptions, that minimization looks like turning down prediction error. Here's the trick: it's not just the states of the world that get optimized—like where the object is and what it is—it's also the precision with which you think those states are encoded.

If precision is itself state-dependent, then attention comes along for free. The brain starts allocating gain to the channels it believes are most reliable, and all of a sudden you see hallmark attention effects—faster detection at cued locations and suppression of distractors—those early and late blips in event-related potentials emerge from the same inferential machinery.

Under the hood sits a hierarchical generative model. Causes live at higher levels, states live below them, and the two together generate the sensory stream through a set of nonlinear mappings. The brain keeps a running belief about these quantities—statisticians would call it a recognition density—and a standard move is to approximate that belief as Gaussian, with a mean and a covariance.

Flip the covariance and you get precision, and that precision shows up everywhere: it scales how strongly prediction errors pull on your beliefs. There's a compact statement of this idea in the math. The free energy, when you assume a Gaussian belief, is dominated by a simple quadratic in the prediction errors for data, for how fast states are changing, and for parameters; the curvature of that quadratic sets the precision.

Match your precision to that curvature and, conveniently, you can update the mean as if it were the only thing that mattered.

Now, not all noise is created equal. In these models, the amplitude of random fluctuations—the sensory noise—depends on context. The log-precision of that noise is a function of hidden states and higher-level causes.

In plainer words, what you think is going on in the scene controls how noisy you expect each channel to be. Those expectations modulate the gain of the forward-going error units, the ones often linked to superficial pyramidal cells that show up in electroencephalography. Slower, neuromodulatory parameters tune this precision too; think of cholinergic systems adjusting the overall gain landscape.

Learning those parameters uses the same gradient rules as learning the content of the generative model, so attention and learning aren't separate hacks—they're two faces of the same updating scheme.

The day-to-day dynamics look like message passing. Bottom-up and lateral error signals tug beliefs toward better explanations, while top-down predictions constantly reshape what counts as a good prior. The algorithmic gloss is a kind of generalized filtering: linearize locally, use the Jacobian to steer your updates, and let curvature terms supply the precision that scales each error.

It's mathy, but the behavior is intuitive. If a channel is precise, its errors count for more and your beliefs move quickly in response. If it's noisy, you ignore it. That's attention, rendered as Bayes.

Bring this down to a classic task: Posner cueing. You get a cue that's valid most of the time—about four out of five in the foundational studies—and then a target flashes left or right. Valid cues speed you up; invalid ones slow you down.

In the precision-weighting story, the cue is nothing mystical. It just sets up a high-precision context for the cued location. When the target arrives, the prediction errors from that channel are amplified and push beliefs fast.

Errors from the uncued side are down-weighted and drift. The precision context itself has inertia, so some of the bias persists after the cue, then decays.

To show this in action, Friston's group built a minimal two-level model. The first level encodes visual inputs. The second level encodes the causes: one for a left target, one for a right target, and one for the central cue.

Hidden states at the first level implement a complementary precision mechanism, like a seesaw—if the left gets high precision, the right gives some up. The sensory noise is modeled with a log-precision that depends on these states. If the signal is strong, the model expects less noise, consistent with Weber's law.

There's also a neat link to Poisson-like sensory variability: represent the signal as the log of a rate, and the precision becomes proportional to that rate. The cue itself is a brief Gaussian bump in time, on the order of a few dozen milliseconds, while the target is a slightly longer biphasic pulse. Hidden states accumulate and decay over hundreds of milliseconds, so the attentional set lingers.

What falls out matches the psychology. In valid trials, the target's representation rides the high-precision channel, and the posterior confidence that a target is present rises quickly; in invalid trials, it ambles. If you take a decision threshold—say you'll respond when posterior confidence hits a set point—you get a reaction time advantage for valid cues of about twenty milliseconds in the model, which is right in the empirical ballpark.

The time course tracks intuition too. Cue-driven effects bloom after a couple of hundred milliseconds and then relax as the precision context fades. Neutral cues, which crank up alertness without biasing space, yield a small timing benefit without favoring one side.

The electrophysiology lines up as well. Simulated event-related potentials show two distinct epochs. Early, valid cues yield bigger, sharper sensory responses—what people label P1 or N1 in the waveform—consistent with a gain boost on sensory prediction errors.

Later, invalid trials kick up larger errors at the level of context, which looks like the brain updating its model of where precision should be after being surprised—think of the P3. Mangun and Hillyard saw this split in real data decades ago, and the match to that pattern is a strong check that the mechanism is pointing the right way. The model also produces slow, anticipatory baseline shifts—contingent negative variation—that reflect the inferred precision in the attended channel before the target even appears.

Put two targets on the screen at once and ask the system to pick. In the classic biased competition view from Desimone and Duncan, stimuli that share a receptive field fight it out, with mutual suppression that grows as you move up the visual hierarchy and receptive fields widen. Imaging in humans shows the same logic in populations: further along the ventral stream, the farther apart two items can be and still suppress each other.

Top-down goal signals and bottom-up salience both tilt the contest. It's a tidy, well-supported story.

What the precision view does is, it doesn't throw that story out; it explains it. Only one precision context can be in force at a time. If the system infers that left is precise now, then left-channel errors get cranked up in their influence and right-channel errors lose steam.

When you present two targets simultaneously, the attended one elicits a strong, narrow representation of its cause. The unattended one is scrubbed down, often to roughly half the response of the attended item. Add a valid cue and the unattended channel's expectation drops further—on the order of twenty percent—with wider uncertainty, because the precision context has shifted even more decisively.

No explicit lateral inhibition needs to be wired into the generative model; the competition comes for free from inferred precision and the nonlinearity that precision injects into the message passing.

Behaviorally, that precision reallocation shows up as the standard speed and accuracy trade-offs. Valid cues buy you speed without much accuracy loss up to a point; invalid cues make you cautious or late, depending on where you set your threshold. In electrophysiology, you get the early sensory boost when your model's precision matches the target, and the late context update when it doesn't.

Luck and colleagues saw the early sensory enhancement track attention in human recordings. The later, larger responses to invalid events look like a context surprise—the brain realizing its precision bet was off and recalibrating.

There's neurobiology ready to plug into these dials. Gamma band synchronization—the fast rhythms between thirty and one hundred hertz—has long been tied to effective gain on synaptic inputs. Cholinergic systems modulate postsynaptic gain and adaptation and are prime candidates for implementing precision control.

Put those facts next to the model's claim that attention is postsynaptic gain on prediction error units, and you have a plausible mechanism bridging computation and tissue. Baseline shifts, early and late event-related components, gamma rhythms, cholinergic tone—they all fall into place as ways the brain encodes, exploits, and updates precision.

If you like equations stated in words, here's the essential one the simulations lean on. Sensory noise is Gaussian with a variance that is the exponential of the negative log-precision. The log-precision itself depends on hidden states and a parameter that sets how strongly context moves precision around.

Because the model represents the sensory signal as a log rate, the precision of that signal equals the underlying rate. That simple relationship makes precision an explicit, movable resource you can reallocate on the fly as the context changes.

There are honest limitations. The model is minimal: two locations, two levels, no explicit lateral wiring, and no attempt to spell out the spatial textures of receptive fields across the cortex. All the attentional boosting appears only when you invert the model—when you run the inference—not because you hard coded gain into the sensory inputs.

That's a feature philosophically, but it leaves open questions about how best to map the abstract nodes onto real circuits. The slower learning of precision over longer timescales is sketched but not yet fit to data, and there's a lot of room for formal comparison to competing models on real datasets.

Even with those caveats, the through-line is compelling. Start with a brain trying to minimize surprise by explaining its inputs. Let that brain learn not just what's out there but how reliable each channel is in the current context.

Then watch attention fall out of the math: valid cue benefits of around twenty milliseconds, early sensory amplification when your precision bet is right, late context updates when it's wrong, and mutual suppression between neighbors when you cram the scene. Desimone and Duncan's competition isn't replaced; it's recast as the inevitable consequence of reallocating precision in a hierarchical inference engine.

Where does this go next? The predictions are specific enough to test. Tie reaction time curves and event-related potentials to model fits, not just qualitative matches.

Manipulate gamma rhythms or cholinergic tone and watch the inferred precision parameters move. Fit the precision learning knobs on timescales of a few hundred trials and see if the model tracks human adjustment to changing cue validities. If attention is the brain's bet on which errors to trust, then every cue, every distractor, and every burst of acetylcholine is a nudge on the odds.

And the free energy view gives us the book the brain is keeping as it lays those bets.

More in Neuroscience