A Hierarchy of Time-Scales and the Brain

Stefan J. Kiebel, Jean Daunizeau, Karl FristonView original
OverviewBalancedharper voice
Imagine you're listening to a bird sing. A single phrase, maybe half a second long, rising then falling. Now ask what it actually takes to hear that as a song rather than noise. You have to track the acoustic texture of each syllable — that's happening in milliseconds. You have to hold the sequence of syllables together — that's seconds. And you have to understand that this phrase belongs to a longer pattern, a context, something slow and structural that the whole melody unfolds inside. One brain, doing all of that at once, across timescales that span orders of magnitude. That is the problem Kiebel, Daunizeau, and Friston set out to solve. Their two thousand eight paper opens with a deceptively simple observation: the environment changes at many speeds simultaneously. A face moves in milliseconds. A mood shifts over minutes. A season changes over months. Speech alone decomposes into nested timescales—instantaneous acoustics, phonemes, words, syntax, pragmatics—each layer riding on top of the one below. The brain must track all of these at once, and it must do so with a single architecture. How? The framework they propose is built on the free-energy principle. The core idea is that adaptive agents — like brains — act to minimize surprise about their sensory input. Surprise here has a precise meaning: it's quantified as two times the natural logarithm of the probability of the sensory data under the agent's model. Minimizing surprise directly is impossible because the agent doesn't have full knowledge of the world. So the agent instead minimizes free-energy, which the paper describes as an upper bound on surprise. The practical consequence is that the brain must maintain internal, dynamical models of environmental causes — models that predict incoming signals and reduce high-dimensional sensation to a handful of latent causes. And crucially, those internal models have to respect the temporal structure of what they're modeling. This is where the mathematics gets interesting. Kiebel and colleagues introduce what they call generalized coordinates of motion. The idea is that the brain doesn't just represent an instantaneous snapshot of the world. Instead, it encodes a local trajectory — the current state, its velocity, its acceleration, and higher derivatives — all at once. Think of it as a Taylor expansion: by specifying enough terms, you capture not just where something is now but where it has been and where it's going. In their simulations, they used six high-order temporal derivatives for hidden states, which gives the model a rich, temporally extended picture of the cause it's trying to track. Hierarchical inference is then built on top of this representation. The model has multiple levels, and the key constraint is that slower, higher levels act as priors on faster, lower levels. Slow dynamics enter as control parameters of fast dynamics. The two levels evolve at different rates. In the birdsong simulation, the fast level had a time constant of zero point two five seconds, and the slow level ran at two seconds, an order of magnitude slower. Predictions flow downward from slow to fast; prediction errors — the mismatch between what was expected and what arrived — flow upward. This lets slow contextual knowledge shape fast perception without the two levels collapsing into a single undifferentiated process. The birdsong simulation is where this becomes vivid and testable. The team built a two-level generative model. At the fast level, a Lorenz attractor—those famous three coupled nonlinear equations that produce a chaotic, butterfly-shaped trajectory—generates the acoustic detail of each syllable. The attractor's outputs drive a simulated sonogram, with amplitude from one output and frequency from another, spanning two to five kilohertz. At the slow level, a second Lorenz attractor runs eight times more slowly and acts as a composer: its output modulates the Rayleigh number — a control parameter — of the fast attractor, thereby choosing which syllable sequence gets produced. The slow system doesn't sing. It decides what gets sung. Perception in the simulation is model inversion: the listening agent uses an online variational scheme to estimate hidden and causal states at both levels from the raw sonogram. Two diagnostic tests reveal what the slow contextual level actually does. First, when the song was interrupted at one point four seconds — the last two syllables removed — the fast level's estimate fell to zero only after about one hundred milliseconds, while the slow level continued along its expected trajectory and produced a large, sustained prediction error. The slow context held on, even after the input stopped. Second, when the slow level was removed entirely and a single-level model was inverted on the same data, the response to the interruption was much faster — on the order of ten milliseconds — and the model accumulated greater prediction error overall when trying to explain normal, uninterrupted song. Under added noise, the two-level model still explained the data reasonably, missing only one syllable, while the single-level model failed entirely. The slow Lorenz attractor isn't decorative. It's what makes the model work. Now comes the anatomical argument, and this is where the paper makes its most ambitious claim. Kiebel, Daunizeau, and Friston argue that the brain's physical layout — from primary sensory areas at the back to prefrontal cortex at the front — is a map of timescales. Caudal, sensory-proximate areas represent fast, short-lived causes. Progressively more rostral regions represent slower, more persistent aspects of the environment. The paper organizes this into a rough table: sensory and association cortex operates over milliseconds to hundreds of milliseconds; primary motor and premotor cortex over tens of milliseconds to seconds; lateral prefrontal cortex over tens of seconds to much longer; and orbitofrontal cortex represents the most temporally stable environmental states over very long periods. The gradient is continuous, and it runs front to back. The anatomical evidence for this lives in the distinction between forward and backward cortical connections. Forward connections — carrying sensory evidence up the hierarchy — arise from superficial pyramidal cells and terminate in layer four of higher areas. They are fast and driving. Backward connections run the other way: they arise from deep pyramidal cells, target supra and infra-granular layers of lower areas, and are associated with slower synaptic time constants and nonlinear N-methyl-D-aspartate receptors. They are modulatory. This laminar and synaptic asymmetry is exactly what the model predicts: fast driving evidence travels upward, while slow contextual priors travel downward. The one hundred millisecond versus ten millisecond comparison from the birdsong simulation isn't just a computational result. It's a prediction about the temporal signature of backward connections in real cortex. Seen through this lens, the prefrontal cortex gets reframed. It isn't simply the seat of abstract reasoning. It's the cortical territory that represents the slowest changing aspects of the environment — the contextual backdrop against which everything faster plays out. Damage to prefrontal cortex disrupts planning and context sensitivity not because reasoning per se is broken, but because the slow representations that constrain fast processing are gone. The frontal lobe is doing what the slow Lorenz attractor did in the simulation: holding the contextual structure that organizes everything beneath it. The framework makes concrete, testable predictions. Manipulating the timescale of sensory input — speeding up or slowing down birdsong, for instance — should selectively engage different cortical levels. Areas close to primary sensory cortex should respond to fast manipulations; more rostral areas should respond to slow ones. The model also predicts that the specific cortical area engaged by a task can be inferred from where that task sits along the rostro-caudal gradient — which reflects anatomical distance from primary sensory areas and, by the argument of this paper, corresponds directly to the timescale of the environmental cause being represented. What makes this framework compelling is its parsimony. A single principle — minimize free-energy using a hierarchical model with separated timescales — explains why the brain is organized the way it is, why prefrontal lesions have the effects they do, why backward connections are slower and more modulatory than forward ones, and why a birdsong simulation with a slow Lorenz attractor outperforms one without. The world has time within time. And the brain, Kiebel, Daunizeau, and Friston argue, is built to match it — not by accident, but because that matching is exactly what it means to model the world well. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Imagine you're listening to a bird sing. A single phrase, maybe half a second long, rising then falling. Now ask what it actually takes to hear that as a song rather than noise. You have to track the acoustic texture of each syllable — that's happening in milliseconds. You have to hold the sequence of syllables together — that's seconds. And you have to understand that this phrase belongs to a longer pattern, a context, something slow and structural that the whole melody unfolds inside. One brain, doing all of that at once, across timescales that span orders of magnitude. That is the problem Kiebel, Daunizeau, and Friston set out to solve. Their two thousand eight paper opens with a deceptively simple observation: the environment changes at many speeds simultaneously. A face moves in milliseconds. A mood shifts over minutes. A season changes over months. Speech alone decomposes into nested timescales—instantaneous acoustics, phonemes, words, syntax, pragmatics—each layer riding on top of the one below. The brain must track all of these at once, and it must do so with a single architecture. How? The framework they propose is built on the free-energy principle. The core idea is that adaptive agents — like brains — act to minimize surprise about their sensory input. Surprise here has a precise meaning: it's quantified as two times the natural logarithm of the probability of the sensory data under the agent's model.

Minimizing surprise directly is impossible because the agent doesn't have full knowledge of the world. So the agent instead minimizes free-energy, which the paper describes as an upper bound on surprise. The practical consequence is that the brain must maintain internal, dynamical models of environmental causes — models that predict incoming signals and reduce high-dimensional sensation to a handful of latent causes. And crucially, those internal models have to respect the temporal structure of what they're modeling. This is where the mathematics gets interesting. Kiebel and colleagues introduce what they call generalized coordinates of motion. The idea is that the brain doesn't just represent an instantaneous snapshot of the world. Instead, it encodes a local trajectory — the current state, its velocity, its acceleration, and higher derivatives — all at once. Think of it as a Taylor expansion: by specifying enough terms, you capture not just where something is now but where it has been and where it's going. In their simulations, they used six high-order temporal derivatives for hidden states, which gives the model a rich, temporally extended picture of the cause it's trying to track. Hierarchical inference is then built on top of this representation. The model has multiple levels, and the key constraint is that slower, higher levels act as priors on faster, lower levels. Slow dynamics enter as control parameters of fast dynamics.

The two levels evolve at different rates. In the birdsong simulation, the fast level had a time constant of zero point two five seconds, and the slow level ran at two seconds, an order of magnitude slower. Predictions flow downward from slow to fast; prediction errors — the mismatch between what was expected and what arrived — flow upward. This lets slow contextual knowledge shape fast perception without the two levels collapsing into a single undifferentiated process. The birdsong simulation is where this becomes vivid and testable. The team built a two-level generative model. At the fast level, a Lorenz attractor—those famous three coupled nonlinear equations that produce a chaotic, butterfly-shaped trajectory—generates the acoustic detail of each syllable. The attractor's outputs drive a simulated sonogram, with amplitude from one output and frequency from another, spanning two to five kilohertz. At the slow level, a second Lorenz attractor runs eight times more slowly and acts as a composer: its output modulates the Rayleigh number — a control parameter — of the fast attractor, thereby choosing which syllable sequence gets produced. The slow system doesn't sing. It decides what gets sung.

Perception in the simulation is model inversion: the listening agent uses an online variational scheme to estimate hidden and causal states at both levels from the raw sonogram. Two diagnostic tests reveal what the slow contextual level actually does. First, when the song was interrupted at one point four seconds — the last two syllables removed — the fast level's estimate fell to zero only after about one hundred milliseconds, while the slow level continued along its expected trajectory and produced a large, sustained prediction error. The slow context held on, even after the input stopped. Second, when the slow level was removed entirely and a single-level model was inverted on the same data, the response to the interruption was much faster — on the order of ten milliseconds — and the model accumulated greater prediction error overall when trying to explain normal, uninterrupted song. Under added noise, the two-level model still explained the data reasonably, missing only one syllable, while the single-level model failed entirely. The slow Lorenz attractor isn't decorative. It's what makes the model work. Now comes the anatomical argument, and this is where the paper makes its most ambitious claim. Kiebel, Daunizeau, and Friston argue that the brain's physical layout — from primary sensory areas at the back to prefrontal cortex at the front — is a map of timescales. Caudal, sensory-proximate areas represent fast, short-lived causes.

Progressively more rostral regions represent slower, more persistent aspects of the environment. The paper organizes this into a rough table: sensory and association cortex operates over milliseconds to hundreds of milliseconds; primary motor and premotor cortex over tens of milliseconds to seconds; lateral prefrontal cortex over tens of seconds to much longer; and orbitofrontal cortex represents the most temporally stable environmental states over very long periods. The gradient is continuous, and it runs front to back. The anatomical evidence for this lives in the distinction between forward and backward cortical connections. Forward connections — carrying sensory evidence up the hierarchy — arise from superficial pyramidal cells and terminate in layer four of higher areas. They are fast and driving. Backward connections run the other way: they arise from deep pyramidal cells, target supra and infra-granular layers of lower areas, and are associated with slower synaptic time constants and nonlinear N-methyl-D-aspartate receptors. They are modulatory. This laminar and synaptic asymmetry is exactly what the model predicts: fast driving evidence travels upward, while slow contextual priors travel downward. The one hundred millisecond versus ten millisecond comparison from the birdsong simulation isn't just a computational result. It's a prediction about the temporal signature of backward connections in real cortex.

Seen through this lens, the prefrontal cortex gets reframed. It isn't simply the seat of abstract reasoning. It's the cortical territory that represents the slowest changing aspects of the environment — the contextual backdrop against which everything faster plays out. Damage to prefrontal cortex disrupts planning and context sensitivity not because reasoning per se is broken, but because the slow representations that constrain fast processing are gone. The frontal lobe is doing what the slow Lorenz attractor did in the simulation: holding the contextual structure that organizes everything beneath it. The framework makes concrete, testable predictions. Manipulating the timescale of sensory input — speeding up or slowing down birdsong, for instance — should selectively engage different cortical levels. Areas close to primary sensory cortex should respond to fast manipulations; more rostral areas should respond to slow ones. The model also predicts that the specific cortical area engaged by a task can be inferred from where that task sits along the rostro-caudal gradient — which reflects anatomical distance from primary sensory areas and, by the argument of this paper, corresponds directly to the timescale of the environmental cause being represented.

What makes this framework compelling is its parsimony. A single principle — minimize free-energy using a hierarchical model with separated timescales — explains why the brain is organized the way it is, why prefrontal lesions have the effects they do, why backward connections are slower and more modulatory than forward ones, and why a birdsong simulation with a slow Lorenz attractor outperforms one without. The world has time within time. And the brain, Kiebel, Daunizeau, and Friston argue, is built to match it — not by accident, but because that matching is exactly what it means to model the world well. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Neuroscience