The Goldilocks EffectHuman Infants Allocate Attention to Visual Sequences That Are Neither Too Simple Nor Too Complex

Celeste Kidd, Steven T. Piantadosi, Richard Ν. AslinView original
OverviewBalancedalloy voice
If you've ever watched a baby lock onto something and then, just as suddenly, look away, you've seen a puzzle that has bothered psychologists for decades. Why that moment? Why then? One elegant answer is the Goldilocks idea: not too simple and not too chaotic—infants stick with things that hit an intermediate, learnable groove. In 2012, Celeste Kidd, Steven Piantadosi, and Richard Aslin took that intuition and gave it substance. They argued that babies regulate attention to keep their information intake in the sweet spot. And crucially, they didn't leave it at metaphor. They built a model that predicts, event by event, when a baby will keep watching and when they'll look away. Here's the move. They cast the infant as an "ideal observer" who tracks how likely different events are and updates those beliefs as new evidence comes in. The technical backbone is a Dirichlet-multinomial model—a simple Bayesian learner over a finite menu of events. At any moment, the model assigns a probability to each possible next event, then calls the event's "complexity" the negative log of that probability. In plain terms, the more surprising the event is under the learner's current expectations, the higher its complexity; the more predictable, the lower. Set the prior to be neutral—what statisticians call a uniform Dirichlet with a concentration parameter of one—and you've said, before seeing anything, treat all outcomes as equally likely. In their one-box display, there are two outcomes, so the prior expectation is a fifty-fifty split; in their three-box display, there are three outcomes, so it's one third each. The only equation you need to keep in your head is this: the model's predicted probability for the next event is the expected value of that event's chance under the updated beliefs, and complexity is just the negative log of that number. Low probability means high complexity. High probability means low complexity. Now, how do you link that to a squirmy seven-month-old? By being very concrete about what counts as attention. Kidd and colleagues brought infants into the lab and showed them simple animated scenes with occluding boxes. On each event, a box opened, an object popped out for one second, and then it hid again for one second. The rule was straightforward: the moment an infant looked off-screen for more than one second, the trial ended. If they never looked away, the sequence timed out at one minute. That off-screen definition becomes the dependent measure: the hazard of a look-away at the next event as a function of how surprising that event is to the model. They ran two versions of the task. In Experiment 1, it's a single box. The first reveal always shows an object; after that, the chance that the object appears on any given lift is set for that trial and can be anything from never to always, sampled across twenty-one evenly spaced values. Forty-two infants, on average seven point nine months old, each saw forty-two trials, two at each probability. This design gives you sequences that range from mind-numbingly repetitive to frustratingly capricious, and everything in between. The point is to give the ideal observer a diet of variable predictability so that complexity, defined by that observer, actually wanders across the range. Experiment 2 adds a layer. Now three different boxes share the stage, each hiding its own object. One object pops out per event, then retreats. Because the pattern can bounce among three places, the sequence lives in a richer probabilistic space. Thirty infants, about seven point six months old, watched thirty-two of these multi-object sequences. Objects, boxes, and order were randomized, but the underlying structures—the actual sequences of reveals—were held constant across infants so the model and the behavior could be aligned in a clean way. What came out of Experiment 1 is the picture the Goldilocks metaphor promised: a U-shaped curve. Babies were more likely to look away when the next event was either very predictable or very surprising, and least likely to look away in the middle. A flexible nonparametric fit to the raw data, a generalized additive model with a binomial link, captured this nonlinearity, but the decisive test was a survival analysis. They used a Cox regression, which is the right tool when the clock stops at the moment of a look-away and you want to model the hazard of that event while leaving the baseline time course unconstrained. The squared term for complexity—remember, complexity is negative log probability—was significant. The coefficient on squared complexity was about zero point zero five two, which translates into a roughly five percent increase in the look-away hazard for each standard-deviation step in squared surprisal. And the place where attention held best? Right around one point two five bits of information per event—smack in the intermediate range. Pause on that number for a second. Bits are a way to measure information: one bit is the surprise of a fair coin flip. So one point two five bits is not trivial, not overwhelming. It's a groove. What this means behaviorally is that infants weren't just enthralled by shock and novelty, and they weren't just zoning out to predictable loops either. They were hanging around when the sequence asked them to learn, but not to struggle. Experiment 2 turned the screw by giving infants sequences that invite not just frequency tracking—how often does each object appear?—but dependency tracking—what tends to follow what. If infants are sensitive to the order of events, then a transitional model that pays attention to what just happened should do a better job than one that treats each event as independent. The analysis started with the simpler, non-transitional model and again found the U-shape. The squared complexity term came in at about zero point two six nine in the Cox regression; that's a thirty-one percent bump in the look-away hazard per standard deviation of squared surprisal. There was also a linear term, negative in sign, hinting that the right-hand side of the U—those very surprising moments—drove look-aways a bit more than the very predictable ones. And, as a pragmatic detail, trial number mattered too: a small increase in look-aways across the session, consistent with fatigue. Then they asked the richer question. What if the model takes into account what just happened, so that the expected probability of the next event is conditioned on the transition from the previous one? That Markov-Dirichlet-multinomial variant—same Bayesian heart, now with order—fit the behavior even better. Under this transitional model, the squared complexity effect was stronger, with a coefficient around zero point three five six, which means the hazard ticked up by about forty-three percent per standard deviation of squared surprisal. Trial number still nudged things, and there was a borderline hint that an object's first appearance carried a bit more pull than later ones. But the key test was head-to-head. Put both predictors—transitional and non-transitional complexity—into the same Cox model. Transitional complexity stayed significant, with a coefficient around zero point two eight nine and a p-value below zero point zero one. The non-transitional version all but disappeared, hovering near zero with no statistical weight. That comparison pins the story: infants weren't just tracking how often each thing popped up. They were, at seven to eight months, sensitive to the structure in the stream—the probabilities that knit one moment to the next. Under the hood, the modeling choices were spare and principled. They fixed the Dirichlet prior's concentration at one, a neutral setting that says, before any data, treat all outcomes as equally plausible. The complexity values they fed into the regressions were standardized—mean zero, unit variance—so coefficients are on a common scale across displays and infants. And the survival framework respected the core reality of infant data: once the baby looks away, the rest of the trial is gone, so you shouldn't pretend the time that would have happened after that was observed. A generalized additive model, used as a descriptive check, traced the U-shape without assuming it, which helps separate the visual pattern from the inferential claim. If you're wondering whether boring procedural wrinkles might be driving the effect—eye-tracker hiccups, unusually long trials—they were careful there too. The rule for stopping a trial was fixed. Look away for more than one second, the event sequence ends. Anything that timed out at sixty seconds dropped from analysis, as did trials with too few events to set expectations. Those aren't the headline numbers, but they matter in making sure the U-shape isn't an artifact of noisy tails or technical false stops. Across the two experiments, the infants were full-term, healthy by parental report, and of very similar ages, which keeps the developmental context tight. There's a temptation to hear "Goldilocks" and think it's a cute label hung on a messy dataset. What Kidd, Piantadosi, and Aslin showed is the opposite. They constructed a content-rich metric—negative log probability under a running Bayesian belief state—and found that attention tracks that metric in a principled way. You could write the core relationship as: the probability of a look-away at the next moment depends on the square of the event's surprisal, with the minimum hazard at an intermediate information rate. That's not just a curve on a plot; it's a behavioral signature of a control policy for learning. It also nudges a deeper point about what infants are doing when they watch the world. They're not passively soaking up stimulation until they get overwhelmed. They're regulating. They're holding onto streams that match the bandwidth of their current model and dropping the ones that would either bore them or break them. The transitional-model win in Experiment 2 tightens that picture by showing that even in the first year, babies are keyed to temporal structure—those if-then relationships that define real-world sequences. Of course, attention isn't one dial with a single setting. The authors are explicit about limits. The Goldilocks pattern isn't the only force in play, and it may vary across domains, species, or sensory modalities. These were visual sequences with objects popping out of boxes, not speech or music or social interaction. And the developmental window was narrow: seven to eight months. Still, when you see the same U-shape across two quite different displays, with a preferred center of gravity near one point two five bits and a transition-sensitive model winning out over a frequency-only one, you start to think this is a general strategy, not a lab trick. If you're building a theory of how humans learn so much so fast with such limited resources, this matters. A learner that keeps itself in the learnable zone can extract structure efficiently without getting swamped. That idea shows up in machine learning under names like curiosity and uncertainty sampling; here you're watching it in a nursery chair. The next steps—across language, vision, action—will ask how this control policy interacts with the specific structures of each domain. But the evidence we have is already a payoff: babies modulate attention to manage information intake, and we can predict that modulation with a remarkably simple probabilistic learner. When you picture that infant gaze now—the long stare, the quick flick away—you can imagine the invisible math running in the background. Not heavy calculus. Just a quiet estimate of "What will happen next?" and a decision to keep watching if the answer promises just the right amount of surprise. Not too little. Not too much. Just right.

If you've ever watched a baby lock onto something and then, just as suddenly, look away, you've seen a puzzle that has bothered psychologists for decades. Why that moment? Why then?

One elegant answer is the Goldilocks idea: not too simple and not too chaotic—infants stick with things that hit an intermediate, learnable groove. In 2012, Celeste Kidd, Steven Piantadosi, and Richard Aslin took that intuition and gave it substance. They argued that babies regulate attention to keep their information intake in the sweet spot.

And crucially, they didn't leave it at metaphor. They built a model that predicts, event by event, when a baby will keep watching and when they'll look away.

Here's the move. They cast the infant as an "ideal observer" who tracks how likely different events are and updates those beliefs as new evidence comes in. The technical backbone is a Dirichlet-multinomial model—a simple Bayesian learner over a finite menu of events.

At any moment, the model assigns a probability to each possible next event, then calls the event's "complexity" the negative log of that probability. In plain terms, the more surprising the event is under the learner's current expectations, the higher its complexity; the more predictable, the lower. Set the prior to be neutral—what statisticians call a uniform Dirichlet with a concentration parameter of one—and you've said, before seeing anything, treat all outcomes as equally likely.

In their one-box display, there are two outcomes, so the prior expectation is a fifty-fifty split; in their three-box display, there are three outcomes, so it's one third each. The only equation you need to keep in your head is this: the model's predicted probability for the next event is the expected value of that event's chance under the updated beliefs, and complexity is just the negative log of that number. Low probability means high complexity. High probability means low complexity.

Now, how do you link that to a squirmy seven-month-old? By being very concrete about what counts as attention. Kidd and colleagues brought infants into the lab and showed them simple animated scenes with occluding boxes.

On each event, a box opened, an object popped out for one second, and then it hid again for one second. The rule was straightforward: the moment an infant looked off-screen for more than one second, the trial ended. If they never looked away, the sequence timed out at one minute.

That off-screen definition becomes the dependent measure: the hazard of a look-away at the next event as a function of how surprising that event is to the model.

They ran two versions of the task. In Experiment 1, it's a single box. The first reveal always shows an object; after that, the chance that the object appears on any given lift is set for that trial and can be anything from never to always, sampled across twenty-one evenly spaced values.

Forty-two infants, on average seven point nine months old, each saw forty-two trials, two at each probability. This design gives you sequences that range from mind-numbingly repetitive to frustratingly capricious, and everything in between. The point is to give the ideal observer a diet of variable predictability so that complexity, defined by that observer, actually wanders across the range.

Experiment 2 adds a layer. Now three different boxes share the stage, each hiding its own object. One object pops out per event, then retreats.

Because the pattern can bounce among three places, the sequence lives in a richer probabilistic space. Thirty infants, about seven point six months old, watched thirty-two of these multi-object sequences. Objects, boxes, and order were randomized, but the underlying structures—the actual sequences of reveals—were held constant across infants so the model and the behavior could be aligned in a clean way.

What came out of Experiment 1 is the picture the Goldilocks metaphor promised: a U-shaped curve. Babies were more likely to look away when the next event was either very predictable or very surprising, and least likely to look away in the middle. A flexible nonparametric fit to the raw data, a generalized additive model with a binomial link, captured this nonlinearity, but the decisive test was a survival analysis.

They used a Cox regression, which is the right tool when the clock stops at the moment of a look-away and you want to model the hazard of that event while leaving the baseline time course unconstrained. The squared term for complexity—remember, complexity is negative log probability—was significant. The coefficient on squared complexity was about zero point zero five two, which translates into a roughly five percent increase in the look-away hazard for each standard-deviation step in squared surprisal.

And the place where attention held best? Right around one point two five bits of information per event—smack in the intermediate range.

Pause on that number for a second. Bits are a way to measure information: one bit is the surprise of a fair coin flip. So one point two five bits is not trivial, not overwhelming.

It's a groove. What this means behaviorally is that infants weren't just enthralled by shock and novelty, and they weren't just zoning out to predictable loops either. They were hanging around when the sequence asked them to learn, but not to struggle.

Experiment 2 turned the screw by giving infants sequences that invite not just frequency tracking—how often does each object appear?—but dependency tracking—what tends to follow what. If infants are sensitive to the order of events, then a transitional model that pays attention to what just happened should do a better job than one that treats each event as independent. The analysis started with the simpler, non-transitional model and again found the U-shape.

The squared complexity term came in at about zero point two six nine in the Cox regression; that's a thirty-one percent bump in the look-away hazard per standard deviation of squared surprisal. There was also a linear term, negative in sign, hinting that the right-hand side of the U—those very surprising moments—drove look-aways a bit more than the very predictable ones. And, as a pragmatic detail, trial number mattered too: a small increase in look-aways across the session, consistent with fatigue.

Then they asked the richer question. What if the model takes into account what just happened, so that the expected probability of the next event is conditioned on the transition from the previous one? That Markov-Dirichlet-multinomial variant—same Bayesian heart, now with order—fit the behavior even better.

Under this transitional model, the squared complexity effect was stronger, with a coefficient around zero point three five six, which means the hazard ticked up by about forty-three percent per standard deviation of squared surprisal. Trial number still nudged things, and there was a borderline hint that an object's first appearance carried a bit more pull than later ones. But the key test was head-to-head.

Put both predictors—transitional and non-transitional complexity—into the same Cox model. Transitional complexity stayed significant, with a coefficient around zero point two eight nine and a p-value below zero point zero one. The non-transitional version all but disappeared, hovering near zero with no statistical weight.

That comparison pins the story: infants weren't just tracking how often each thing popped up. They were, at seven to eight months, sensitive to the structure in the stream—the probabilities that knit one moment to the next.

Under the hood, the modeling choices were spare and principled. They fixed the Dirichlet prior's concentration at one, a neutral setting that says, before any data, treat all outcomes as equally plausible. The complexity values they fed into the regressions were standardized—mean zero, unit variance—so coefficients are on a common scale across displays and infants.

And the survival framework respected the core reality of infant data: once the baby looks away, the rest of the trial is gone, so you shouldn't pretend the time that would have happened after that was observed. A generalized additive model, used as a descriptive check, traced the U-shape without assuming it, which helps separate the visual pattern from the inferential claim.

If you're wondering whether boring procedural wrinkles might be driving the effect—eye-tracker hiccups, unusually long trials—they were careful there too. The rule for stopping a trial was fixed. Look away for more than one second, the event sequence ends.

Anything that timed out at sixty seconds dropped from analysis, as did trials with too few events to set expectations. Those aren't the headline numbers, but they matter in making sure the U-shape isn't an artifact of noisy tails or technical false stops. Across the two experiments, the infants were full-term, healthy by parental report, and of very similar ages, which keeps the developmental context tight.

There's a temptation to hear "Goldilocks" and think it's a cute label hung on a messy dataset. What Kidd, Piantadosi, and Aslin showed is the opposite. They constructed a content-rich metric—negative log probability under a running Bayesian belief state—and found that attention tracks that metric in a principled way.

You could write the core relationship as: the probability of a look-away at the next moment depends on the square of the event's surprisal, with the minimum hazard at an intermediate information rate. That's not just a curve on a plot; it's a behavioral signature of a control policy for learning.

It also nudges a deeper point about what infants are doing when they watch the world. They're not passively soaking up stimulation until they get overwhelmed. They're regulating.

They're holding onto streams that match the bandwidth of their current model and dropping the ones that would either bore them or break them. The transitional-model win in Experiment 2 tightens that picture by showing that even in the first year, babies are keyed to temporal structure—those if-then relationships that define real-world sequences.

Of course, attention isn't one dial with a single setting. The authors are explicit about limits. The Goldilocks pattern isn't the only force in play, and it may vary across domains, species, or sensory modalities.

These were visual sequences with objects popping out of boxes, not speech or music or social interaction. And the developmental window was narrow: seven to eight months. Still, when you see the same U-shape across two quite different displays, with a preferred center of gravity near one point two five bits and a transition-sensitive model winning out over a frequency-only one, you start to think this is a general strategy, not a lab trick.

If you're building a theory of how humans learn so much so fast with such limited resources, this matters. A learner that keeps itself in the learnable zone can extract structure efficiently without getting swamped. That idea shows up in machine learning under names like curiosity and uncertainty sampling; here you're watching it in a nursery chair.

The next steps—across language, vision, action—will ask how this control policy interacts with the specific structures of each domain. But the evidence we have is already a payoff: babies modulate attention to manage information intake, and we can predict that modulation with a remarkably simple probabilistic learner.

When you picture that infant gaze now—the long stare, the quick flick away—you can imagine the invisible math running in the background. Not heavy calculus. Just a quiet estimate of "What will happen next?" and a decision to keep watching if the answer promises just the right amount of surprise. Not too little. Not too much. Just right.

More in Psychology