Embodied Cognition is Not What you Think it is

Andrew D. Wilson, Sabrina GolonkaView original
OverviewBalancedhelen voice
The field of embodied cognition has produced hundreds of studies showing that body states change mental states. For example, if you hold a warm cup of coffee, you might rate strangers as warmer people. If you clench a fist, you may feel more in control. That sounds radical. Andrew Wilson and Sabrina Golonka argue it's actually the most conservative version of the idea, and that the genuinely radical version has barely been tested at all. Here's the distinction they draw. The common definition of embodied cognition suggests that the body modifies mental states. But that framing still assumes the brain contains a mental state, and the body just tweaks it. The stronger claim is that if cognition can span brain, body, and environment, then those brain-based mental states might not exist to be modified in the first place. Cognition could be a distributed system assembled from resources drawn from the brain, the body, and the world, with no central internal representation anywhere in sight. Wilson and Golonka call this the replacement hypothesis — not supplement but replace. To test that hypothesis seriously, they propose a four-step research program. Step one: task analysis. Characterize the problem from the agent's first-person perspective, rather than the experimenter's bird's-eye view. Organisms often use smart, task-specific strategies instead of general-purpose algorithms, and you can only see those strategies if you start from the right vantage point. Step two: identify all task-relevant resources available to the agent, including brain, body, environment, and the relationships among them. Step three: specify how those resources can be assembled into a system that actually solves the task, typically a dynamical system distributed across all three. Step four: test whether the agent actually uses that assembly, usually through perturbation experiments, because a system built from specific components responds to perturbations in ways that reveal its structure. This is not a rebranding of standard cognitive science. It's a different set of questions aimed at a different kind of answer. The outfielder problem is their first worked example, and it's a good one. Standard cognitive science frames catching a fly ball as an internal physics problem: the fielder perceives the ball's initial direction, speed, and angle, runs a mental simulation of projectile motion, computes where the ball will land, and runs there. Saxberg laid out this account in detail. Wilson finds it unconvincing on practical grounds. At the distances involved, an outfielder might be two hundred fifty feet from home plate, and the optical projection of a baseball is tiny. Cutting and Vishton argued that recovering useful distance estimates from such small optical cues would be highly error-prone. Shaffer and McBeath provided evidence that a simulation-based solution is unlikely in practice. The embodied alternative doesn't predict but controls continuously. The fielder keeps a simple optical variable steady by adjusting their movement in real time. Wilson describes two candidate strategies. Optical Acceleration Cancellation, proposed by Chapman in nineteen sixty-eight and tested by Fink and colleagues, has the fielder align with the ball and run so that the ball appears to move at constant velocity in the visual field. If it still seems to accelerate, that's an error signal to correct. Linear Optical Trajectory, supported by McBeath and colleagues, has the fielder move laterally so the ball appears to trace a straight line rather than a curve. Which strategy applies depends on geometry: Optical Acceleration Cancellation works better when the ball is coming nearly straight at you, while Linear Optical Trajectory works when it's heading to one side. Empirical evidence generally favors Linear Optical Trajectory. Both strategies provide continuous error correction, and both allow the fielder to infer whether a ball is even catchable. If you can't eliminate the apparent curvature by running faster, it isn't. No physics engine required. The coupling between movement and visual information is the solution. The A-not-B task tells a similar story at the other end of development. Piaget's setup is simple: hide a toy at location A several times, let the infant retrieve it, then move the toy to location B in full view. Infants around eight to twelve months keep reaching to location A. The standard explanation invokes a performance-competence gap: the infant has the right concept of object permanence somewhere in their brain, but their reaching behavior fails to access it. The error is viewed as a deficit in execution, not knowledge. Thelen and colleagues rejected that framing entirely. They began, as Wilson and Golonka's framework demands, with a detailed inventory of the task's actual resources, including continuous visual input from two wells, experimenter cues, a short delay, and a reach that must be planned and executed. Then they noticed something important — a long list of task parameters reliably changes the error rate. Distance to targets, distinctiveness of the covers, length of the delay, whether the infant has been moved, and whether the goal is food or a toy all matter. If object knowledge were the operative variable, none of that should be influential. Thelen's formal model implements a reach space where prior activations leave traces, and when input shifts to location B, the field already shaped by reaching to location A can pull the system back to location A. The model reproduces the error at the empirically observed rates. Infants err roughly seventy to eighty percent of the time, depending on developmental status, and it contains no object-concept resource at all. Further, it predicted that the error can occur even when there is no hidden object, and that similar reach biases appear in children up to age eleven and in adults. The A-not-B error stops being a window into object knowledge and instead becomes an emergent outcome of a perceiving-acting system shaped by history, posture, and perceptual context. Robotics supplies the third pillar of evidence, and Wilson and Golonka use it to demonstrate that the replacement hypothesis isn't just philosophically appealing — it produces working systems. Maris and te Boekhorst's Didabots operate on a single rule: turn away from a detected obstacle. With the front sensor deactivated, a robot that bumps into a block head-on just keeps pushing until another sensor fires and triggers a turn. With several robots and several blocks, this simple sensorimotor rule plus body geometry produced apparent tidying — blocks end up in heaps. The tidying isn't in the control program; it emerges from the rule, the environment, and the body's geometry. Passive-dynamic walking goes further. McGeer in nineteen ninety and Collins and colleagues in two thousand five showed that robots with no motors and no onboard algorithms can reproduce human gait simply by being correctly built. The form of walking falls out of body mechanics and gravity interacting in real time. Work at MIT demonstrated that adding simple control algorithms to such a body allows the same simple algorithm to produce a wide variety of locomotion behaviors depending on body configuration alone. The gait is not represented anywhere; it's simply a consequence of the system. Webb's robot crickets provide an ethological example. Female crickets have paired eardrums connected by a tube. Sound entering from the side produces out-of-phase signals that amplify the signal on one side. This drives turning interneurons whose decay rates are tuned to species-typical song frequencies. There is no stored song template or comparison. Webb's robots replicate this architecture and successfully localize sound sources. Hedwig and Webb later observed that real female crickets begin moving before they could have processed a full song, consistent with a chirp-onset strategy, exactly what the model predicts. The behavior works because the body is built the right way, and the environment provides the right information. Language is where Wilson and Golonka push the argument into genuinely hard territory. They apply the same four steps: characterize the first-person task, identify resources, specify the assembly, and test the agent. Linguistic information, they argue, should be treated as a task resource like perceptual information — structure in speech, writing, or gesture that agents learn to detect and use, even though its link to meaning is conventional rather than law-like. They're not satisfied with the landmark embodied-language studies. Glenberg and Kaschak found that participants judged sentence sensibility faster when their response movement matched the sentence's implied direction. Stanfield and Zwaan found faster picture verification when object orientation matched the sentence's implied orientation. Wilson and Golonka acknowledge these results show body-language coupling, but find them incomplete. There is no account of the task's resources, no specification of the information content in the stimuli, and no analysis of response dynamics. The coupling is documented but not explained. That critique lands on the lecture's larger point. Showing that the body and language interact is still working within the conservative definition — body states affecting mental states. The replacement hypothesis asks whether linguistic behavior, much like catching a fly ball or reaching for a hidden toy, might be explained entirely as a task-specific assembly of distributed resources, with no internal symbol manipulation at its core. Wilson and Golonka don't claim to have answered that. What they do assert is that cognitive science has the tools to ask the question properly — and that asking it changes what cognition is, not just how we study it. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

The field of embodied cognition has produced hundreds of studies showing that body states change mental states. For example, if you hold a warm cup of coffee, you might rate strangers as warmer people. If you clench a fist, you may feel more in control. That sounds radical. Andrew Wilson and Sabrina Golonka argue it's actually the most conservative version of the idea, and that the genuinely radical version has barely been tested at all. Here's the distinction they draw. The common definition of embodied cognition suggests that the body modifies mental states. But that framing still assumes the brain contains a mental state, and the body just tweaks it. The stronger claim is that if cognition can span brain, body, and environment, then those brain-based mental states might not exist to be modified in the first place. Cognition could be a distributed system assembled from resources drawn from the brain, the body, and the world, with no central internal representation anywhere in sight. Wilson and Golonka call this the replacement hypothesis — not supplement but replace. To test that hypothesis seriously, they propose a four-step research program. Step one: task analysis. Characterize the problem from the agent's first-person perspective, rather than the experimenter's bird's-eye view.

Organisms often use smart, task-specific strategies instead of general-purpose algorithms, and you can only see those strategies if you start from the right vantage point. Step two: identify all task-relevant resources available to the agent, including brain, body, environment, and the relationships among them. Step three: specify how those resources can be assembled into a system that actually solves the task, typically a dynamical system distributed across all three. Step four: test whether the agent actually uses that assembly, usually through perturbation experiments, because a system built from specific components responds to perturbations in ways that reveal its structure. This is not a rebranding of standard cognitive science. It's a different set of questions aimed at a different kind of answer. The outfielder problem is their first worked example, and it's a good one. Standard cognitive science frames catching a fly ball as an internal physics problem: the fielder perceives the ball's initial direction, speed, and angle, runs a mental simulation of projectile motion, computes where the ball will land, and runs there. Saxberg laid out this account in detail.

Wilson finds it unconvincing on practical grounds. At the distances involved, an outfielder might be two hundred fifty feet from home plate, and the optical projection of a baseball is tiny. Cutting and Vishton argued that recovering useful distance estimates from such small optical cues would be highly error-prone. Shaffer and McBeath provided evidence that a simulation-based solution is unlikely in practice. The embodied alternative doesn't predict but controls continuously. The fielder keeps a simple optical variable steady by adjusting their movement in real time. Wilson describes two candidate strategies. Optical Acceleration Cancellation, proposed by Chapman in nineteen sixty-eight and tested by Fink and colleagues, has the fielder align with the ball and run so that the ball appears to move at constant velocity in the visual field. If it still seems to accelerate, that's an error signal to correct. Linear Optical Trajectory, supported by McBeath and colleagues, has the fielder move laterally so the ball appears to trace a straight line rather than a curve. Which strategy applies depends on geometry: Optical Acceleration Cancellation works better when the ball is coming nearly straight at you, while Linear Optical Trajectory works when it's heading to one side. Empirical evidence generally favors Linear Optical Trajectory. Both strategies provide continuous error correction, and both allow the fielder to infer whether a ball is even catchable.

If you can't eliminate the apparent curvature by running faster, it isn't. No physics engine required. The coupling between movement and visual information is the solution. The A-not-B task tells a similar story at the other end of development. Piaget's setup is simple: hide a toy at location A several times, let the infant retrieve it, then move the toy to location B in full view. Infants around eight to twelve months keep reaching to location A. The standard explanation invokes a performance-competence gap: the infant has the right concept of object permanence somewhere in their brain, but their reaching behavior fails to access it. The error is viewed as a deficit in execution, not knowledge. Thelen and colleagues rejected that framing entirely. They began, as Wilson and Golonka's framework demands, with a detailed inventory of the task's actual resources, including continuous visual input from two wells, experimenter cues, a short delay, and a reach that must be planned and executed. Then they noticed something important — a long list of task parameters reliably changes the error rate.

Distance to targets, distinctiveness of the covers, length of the delay, whether the infant has been moved, and whether the goal is food or a toy all matter. If object knowledge were the operative variable, none of that should be influential. Thelen's formal model implements a reach space where prior activations leave traces, and when input shifts to location B, the field already shaped by reaching to location A can pull the system back to location A. The model reproduces the error at the empirically observed rates. Infants err roughly seventy to eighty percent of the time, depending on developmental status, and it contains no object-concept resource at all. Further, it predicted that the error can occur even when there is no hidden object, and that similar reach biases appear in children up to age eleven and in adults. The A-not-B error stops being a window into object knowledge and instead becomes an emergent outcome of a perceiving-acting system shaped by history, posture, and perceptual context. Robotics supplies the third pillar of evidence, and Wilson and Golonka use it to demonstrate that the replacement hypothesis isn't just philosophically appealing — it produces working systems. Maris and te Boekhorst's Didabots operate on a single rule: turn away from a detected obstacle. With the front sensor deactivated, a robot that bumps into a block head-on just keeps pushing until another sensor fires and triggers a turn.

With several robots and several blocks, this simple sensorimotor rule plus body geometry produced apparent tidying — blocks end up in heaps. The tidying isn't in the control program; it emerges from the rule, the environment, and the body's geometry. Passive-dynamic walking goes further. McGeer in nineteen ninety and Collins and colleagues in two thousand five showed that robots with no motors and no onboard algorithms can reproduce human gait simply by being correctly built. The form of walking falls out of body mechanics and gravity interacting in real time. Work at MIT demonstrated that adding simple control algorithms to such a body allows the same simple algorithm to produce a wide variety of locomotion behaviors depending on body configuration alone. The gait is not represented anywhere; it's simply a consequence of the system. Webb's robot crickets provide an ethological example. Female crickets have paired eardrums connected by a tube. Sound entering from the side produces out-of-phase signals that amplify the signal on one side. This drives turning interneurons whose decay rates are tuned to species-typical song frequencies. There is no stored song template or comparison. Webb's robots replicate this architecture and successfully localize sound sources.

Hedwig and Webb later observed that real female crickets begin moving before they could have processed a full song, consistent with a chirp-onset strategy, exactly what the model predicts. The behavior works because the body is built the right way, and the environment provides the right information. Language is where Wilson and Golonka push the argument into genuinely hard territory. They apply the same four steps: characterize the first-person task, identify resources, specify the assembly, and test the agent. Linguistic information, they argue, should be treated as a task resource like perceptual information — structure in speech, writing, or gesture that agents learn to detect and use, even though its link to meaning is conventional rather than law-like. They're not satisfied with the landmark embodied-language studies. Glenberg and Kaschak found that participants judged sentence sensibility faster when their response movement matched the sentence's implied direction. Stanfield and Zwaan found faster picture verification when object orientation matched the sentence's implied orientation. Wilson and Golonka acknowledge these results show body-language coupling, but find them incomplete. There is no account of the task's resources, no specification of the information content in the stimuli, and no analysis of response dynamics. The coupling is documented but not explained.

That critique lands on the lecture's larger point. Showing that the body and language interact is still working within the conservative definition — body states affecting mental states. The replacement hypothesis asks whether linguistic behavior, much like catching a fly ball or reaching for a hidden toy, might be explained entirely as a task-specific assembly of distributed resources, with no internal symbol manipulation at its core. Wilson and Golonka don't claim to have answered that. What they do assert is that cognitive science has the tools to ask the question properly — and that asking it changes what cognition is, not just how we study it. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Psychology