Causal Inference in Multisensory Perception

Konrad P. Körding, Ulrik Beierholm, Wei Ji, Steven R. Quartz, Joshua B. Tenenbaum, Ladan ShamsView original
OverviewBalancedalloy voice
Imagine you're watching a ventriloquist. The puppet's mouth flaps, the voice comes from the speaker's lungs, and yet—somewhere between your ears and your eyes—your brain says, "One thing made both of those." Most of the time, that's a great assumption. But in the real world, sights and sounds sometimes come from different places. So the trick isn't to fuse everything. It's to decide when to fuse and when to keep things apart. That's the heart of the causal inference view of multisensory perception. Before or as the brain combines cues, it asks a simple but profound question: do these signals share a cause or not? If there's one source, combine them. If there are two, don't. Josh Körding and colleagues took that idea, wrote down a model precise enough to make hard predictions, and put it up against human behavior. Here's the idea in plain terms. On each trial, you receive a visual measurement and an auditory measurement. They're noisy—vision is pretty good, but audition is a bit worse—and they might have come from the same place or different places. The model treats "same cause" and "different causes" as two rival stories about the world. Then it uses Bayes' rule to ask, given what I saw and heard, which story is more likely? Formally, the posterior probability of each story is proportional to two things multiplied: how compatible the data are with that story and how plausible that story was to begin with. You can think of the data compatibility as a likelihood and the prior plausibility as, well, a prior over causal structure. Two pieces make this work. First, a spatial prior that gently nudges perceived locations toward the center of the screen—because in the lab setup, stimuli were more likely near midline. Second, the reliability of each sense. If a cue is more precise, it should get more weight when you fuse. In the common-cause case, the optimal estimate of location is a weighted average of the visual and auditory measurements, and the weights depend on inverse variance—higher precision, higher weight. In the two-cause case, each modality trusts itself. And because the brain mixes these conditional estimates proportionally to how likely "one cause" is versus "two causes," the final percept is nonlinear in a very particular way: as the audiovisual disparity grows, the model automatically transitions from fusion toward segregation. A quick sketch of the moving parts. Körding's team kept the model tight: four core parameters. Visual uncertainty, auditory uncertainty, the width of that central spatial prior, and a prior probability that two cues share a cause—p common. They also included a motor-noise term to acknowledge that mapping an internal estimate onto a button press isn't perfectly precise. Crucially, the math stays tractable because they assume Gaussian noise, so all the relevant probabilities can be computed exactly, not approximated. Now, to the experiments. Picture nineteen college students sitting in front of a monitor, wearing headphones. Every trial, a brief patch appears—35 milliseconds of a Gabor grating in noise—at one of five horizontal positions. At the same time, or alone, a 35-millisecond burst of noise plays through the headphones. The sounds are filtered using each person's head-related transfer function—basically tailoring the audio so that "left" sounds left and "right" sounds right for that person. On every trial, participants make two reports: where did the visual event happen, and where did the sound come from? Across combinations of visual and auditory positions, you get a grid of conditions that span small to large disparities. That is your ventriloquist world, under experimental control. One participant was excluded because their auditory localization was extremely poor—essentially unusable for the purpose of modeling—so eighteen remained for the main analysis. The core dataset, then, is rich: five locations by five locations across modalities, with joint response distributions, not just averages. That design matters because the model doesn't just predict a mean bias; it predicts the full pattern of responses conditioned on disparity and reliability. What did people do? Exactly what the causal story would suggest. When the sound and the flash were close together, the perceived sound got pulled toward vision. As the disparity increased, that pull weakened. You can hear the brain hedging: "Probably the same thing, so let's fuse—unless they're too far apart." Körding's model captured this gracefully. When they fit the four parameters to the joint behavioral data, the model explained almost all of the variance—about ninety-seven percent. Split-half checks were similarly strong: fit on one half of the data and test on the other, and you still get R squared values around 0.96 to 0.98. This wasn't a case of overfitting wiggles in the noise. The fitted parameters tell a story about the senses and expectations in that environment. Vision was tight: visual uncertainty around 2.14 units. Audition was noisier: roughly 9.2 units. The spatial prior, that center bias, had a width around 12.3 units. And people brought a modest bias toward a single source—p common was about 0.28, so not a reflex to fuse, more like a cautious openness to it. A motor-noise term near 2.5 units helped bridge internal estimates to button-press responses. A separate look at unimodal pointing suggested an auditory variance component around 7.6 units, consistent with that noisier auditory channel once you factor motor noise in. Put these together, you have a model that doesn't just say "vision wins." It says when and by how much, and how that shifts as the cues drift apart. Here's the punchline in the math, without the symbols. Under the "one cause" story, your best guess for the location is a compromise: take the visual and auditory measurements and average them with weights set by their precision. Under the "two causes" story, your best guesses are just the raw measurements, each corrected by the central prior. Then, mix those conditional guesses according to how likely "one cause" is given what you just heard and saw. That last step—the mixing—is the source of the nonlinearity in behavior. It's why, as disparity grows, the visual pull on auditory perception fades in a characteristic curve rather than dropping off in a straight line. A good model has to beat live alternatives, not just sound plausible. Körding and colleagues stacked it against two obvious competitors: forced fusion, which always integrates the cues, and full segregation, which never does. Forced fusion did what you’d expect—it over-predicted bias at large disparities and under-predicted the graded transition. Its statistical fit wasn't close, explaining only about fifty-six percent of the variance in the joint response data. Full segregation missed the reliable fusion at small disparities. When the team scored models using information criteria—penalizing extra parameters—the causal-inference account still came out decisively ahead. In other words, even after you correct for the flexibility it buys by having a causal prior, the numbers say it's the better explanation. And the test was not light: the dataset had two hundred fifty joint data points per person, and model comparisons were done with likelihoods that respected the full distribution of responses, not just means. There's another angle that matters for everyday perception: judgments of unity. If you ask people, "Did those come from the same event?" their answers track the very same mechanics that drove the localization estimates. When visual and auditory cues were near each other, people were more likely to say "one cause." When they were far apart, they leaned toward "two causes." The model didn't just fit the localization—it also explained roughly seventy-two percent of the variance in those unity reports. And when you look just at the pattern of bias—how much vision tugged the perceived sound around—the account captured about eighty-seven percent of that variance. That's unusually clean convergence: one framework linking "what do you think happened?" with "where do you think it happened?" A couple of design choices are worth pausing on. In some analyses, unimodal trials were set aside when fitting the core audiovisual data to avoid attentional differences between single-cue and dual-cue conditions contaminating the match. The heavy lifting came from the joint responses to bimodal stimuli—the place where the causal question is asked and answered by the brain. And because the sounds were delivered through individualized head-related filters, auditory precision was about as good as you can get with headphones, which makes the reliability difference between vision and audition a meaningful part of the story, not a design artifact. One subtle prediction of the causal account is counterintuitive: when the brain is pretty sure there are two causes, it can show small "negative" biases—repulsions—because each modality's estimate is pulled toward its own prior rather than toward the other cue. That shows up in the data too, albeit in a limited range. It's a reminder that the framework isn't just a softer form of fusion; it's a principled rule about when not to fuse. How does this all connect to the brain? The authors point to probabilistic population coding as one plausible substrate—neurons representing not just a point estimate, but a whole distribution over possible locations, and pooling them in ways that correspond to Bayesian combination under the common-cause story. Areas like the superior colliculus, where multimodal maps line up, are obvious candidates for parts of the circuit. The paper is careful here: it doesn't claim we've found the exact implementation. But it does show that if you assume the brain reasons about causes and uses reliability-weighted fusion when appropriate, you get observed behavior in a way simple averaging cannot achieve. Step back, and the payoff is bigger than ventriloquism. A lot of perception research has treated cue combination as a fixed rule: take a weighted average, done. Körding's work makes the rule conditional. Weighted averages are what you do when, and only when, your brain has reason to believe there's a single thing out there throwing off multiple signals. That decision—about causal structure—drives the pattern of bias, the dependence on disparity, the unity judgments, and the quantitative goodness-of-fit. And it does so with a compact set of parameters that you can measure, interpret, and compare across observers. There are always caveats. The model assumes Gaussian noise and a simple central prior; real environments can get messier. People's p common might change with context, learning, or instructions. And the brain might approximate rather than compute the exact posterior. But in a tightly controlled setting, with individualized auditory cues and joint reports, a causal-inference observer nails what people do. It fuses when it should, separates when it must, and shades smoothly between those regimes as the evidence demands. If you're curious where this goes next, imagine extending the same logic to time—are two flashes and one beep part of one event or two?—or to more than two senses, or to richer priors learned from natural scenes. The blueprint doesn't change: infer the structure, then combine accordingly. What Körding and colleagues showed is that once you make that first, simple decision explicit, a lot of perception that looked mysterious starts to make sense—and it comes with numbers you can test. That's the sweet spot: story meeting substance, with a brain that binds the world together not by default, but by inference.

Imagine you're watching a ventriloquist. The puppet's mouth flaps, the voice comes from the speaker's lungs, and yet—somewhere between your ears and your eyes—your brain says, "One thing made both of those." Most of the time, that's a great assumption. But in the real world, sights and sounds sometimes come from different places.

So the trick isn't to fuse everything. It's to decide when to fuse and when to keep things apart.

That's the heart of the causal inference view of multisensory perception. Before or as the brain combines cues, it asks a simple but profound question: do these signals share a cause or not? If there's one source, combine them.

If there are two, don't. Josh Körding and colleagues took that idea, wrote down a model precise enough to make hard predictions, and put it up against human behavior.

Here's the idea in plain terms. On each trial, you receive a visual measurement and an auditory measurement. They're noisy—vision is pretty good, but audition is a bit worse—and they might have come from the same place or different places.

The model treats "same cause" and "different causes" as two rival stories about the world. Then it uses Bayes' rule to ask, given what I saw and heard, which story is more likely? Formally, the posterior probability of each story is proportional to two things multiplied: how compatible the data are with that story and how plausible that story was to begin with.

You can think of the data compatibility as a likelihood and the prior plausibility as, well, a prior over causal structure.

Two pieces make this work. First, a spatial prior that gently nudges perceived locations toward the center of the screen—because in the lab setup, stimuli were more likely near midline. Second, the reliability of each sense.

If a cue is more precise, it should get more weight when you fuse. In the common-cause case, the optimal estimate of location is a weighted average of the visual and auditory measurements, and the weights depend on inverse variance—higher precision, higher weight. In the two-cause case, each modality trusts itself.

And because the brain mixes these conditional estimates proportionally to how likely "one cause" is versus "two causes," the final percept is nonlinear in a very particular way: as the audiovisual disparity grows, the model automatically transitions from fusion toward segregation.

A quick sketch of the moving parts. Körding's team kept the model tight: four core parameters. Visual uncertainty, auditory uncertainty, the width of that central spatial prior, and a prior probability that two cues share a cause—p common.

They also included a motor-noise term to acknowledge that mapping an internal estimate onto a button press isn't perfectly precise. Crucially, the math stays tractable because they assume Gaussian noise, so all the relevant probabilities can be computed exactly, not approximated.

Now, to the experiments. Picture nineteen college students sitting in front of a monitor, wearing headphones. Every trial, a brief patch appears—35 milliseconds of a Gabor grating in noise—at one of five horizontal positions.

At the same time, or alone, a 35-millisecond burst of noise plays through the headphones. The sounds are filtered using each person's head-related transfer function—basically tailoring the audio so that "left" sounds left and "right" sounds right for that person. On every trial, participants make two reports: where did the visual event happen, and where did the sound come from?

Across combinations of visual and auditory positions, you get a grid of conditions that span small to large disparities. That is your ventriloquist world, under experimental control.

One participant was excluded because their auditory localization was extremely poor—essentially unusable for the purpose of modeling—so eighteen remained for the main analysis. The core dataset, then, is rich: five locations by five locations across modalities, with joint response distributions, not just averages. That design matters because the model doesn't just predict a mean bias; it predicts the full pattern of responses conditioned on disparity and reliability.

What did people do? Exactly what the causal story would suggest. When the sound and the flash were close together, the perceived sound got pulled toward vision.

As the disparity increased, that pull weakened. You can hear the brain hedging: "Probably the same thing, so let's fuse—unless they're too far apart." Körding's model captured this gracefully. When they fit the four parameters to the joint behavioral data, the model explained almost all of the variance—about ninety-seven percent.

Split-half checks were similarly strong: fit on one half of the data and test on the other, and you still get R squared values around 0.96 to 0.98. This wasn't a case of overfitting wiggles in the noise.

The fitted parameters tell a story about the senses and expectations in that environment. Vision was tight: visual uncertainty around 2.14 units. Audition was noisier: roughly 9.2 units.

The spatial prior, that center bias, had a width around 12.3 units. And people brought a modest bias toward a single source—p common was about 0.28, so not a reflex to fuse, more like a cautious openness to it. A motor-noise term near 2.5 units helped bridge internal estimates to button-press responses.

A separate look at unimodal pointing suggested an auditory variance component around 7.6 units, consistent with that noisier auditory channel once you factor motor noise in. Put these together, you have a model that doesn't just say "vision wins." It says when and by how much, and how that shifts as the cues drift apart.

Here's the punchline in the math, without the symbols. Under the "one cause" story, your best guess for the location is a compromise: take the visual and auditory measurements and average them with weights set by their precision. Under the "two causes" story, your best guesses are just the raw measurements, each corrected by the central prior.

Then, mix those conditional guesses according to how likely "one cause" is given what you just heard and saw. That last step—the mixing—is the source of the nonlinearity in behavior. It's why, as disparity grows, the visual pull on auditory perception fades in a characteristic curve rather than dropping off in a straight line.

A good model has to beat live alternatives, not just sound plausible. Körding and colleagues stacked it against two obvious competitors: forced fusion, which always integrates the cues, and full segregation, which never does. Forced fusion did what you’d expect—it over-predicted bias at large disparities and under-predicted the graded transition.

Its statistical fit wasn't close, explaining only about fifty-six percent of the variance in the joint response data. Full segregation missed the reliable fusion at small disparities. When the team scored models using information criteria—penalizing extra parameters—the causal-inference account still came out decisively ahead.

In other words, even after you correct for the flexibility it buys by having a causal prior, the numbers say it's the better explanation. And the test was not light: the dataset had two hundred fifty joint data points per person, and model comparisons were done with likelihoods that respected the full distribution of responses, not just means.

There's another angle that matters for everyday perception: judgments of unity. If you ask people, "Did those come from the same event?" their answers track the very same mechanics that drove the localization estimates. When visual and auditory cues were near each other, people were more likely to say "one cause." When they were far apart, they leaned toward "two causes." The model didn't just fit the localization—it also explained roughly seventy-two percent of the variance in those unity reports.

And when you look just at the pattern of bias—how much vision tugged the perceived sound around—the account captured about eighty-seven percent of that variance. That's unusually clean convergence: one framework linking "what do you think happened?" with "where do you think it happened?"

A couple of design choices are worth pausing on. In some analyses, unimodal trials were set aside when fitting the core audiovisual data to avoid attentional differences between single-cue and dual-cue conditions contaminating the match. The heavy lifting came from the joint responses to bimodal stimuli—the place where the causal question is asked and answered by the brain.

And because the sounds were delivered through individualized head-related filters, auditory precision was about as good as you can get with headphones, which makes the reliability difference between vision and audition a meaningful part of the story, not a design artifact.

One subtle prediction of the causal account is counterintuitive: when the brain is pretty sure there are two causes, it can show small "negative" biases—repulsions—because each modality's estimate is pulled toward its own prior rather than toward the other cue. That shows up in the data too, albeit in a limited range. It's a reminder that the framework isn't just a softer form of fusion; it's a principled rule about when not to fuse.

How does this all connect to the brain? The authors point to probabilistic population coding as one plausible substrate—neurons representing not just a point estimate, but a whole distribution over possible locations, and pooling them in ways that correspond to Bayesian combination under the common-cause story. Areas like the superior colliculus, where multimodal maps line up, are obvious candidates for parts of the circuit.

The paper is careful here: it doesn't claim we've found the exact implementation. But it does show that if you assume the brain reasons about causes and uses reliability-weighted fusion when appropriate, you get observed behavior in a way simple averaging cannot achieve.

Step back, and the payoff is bigger than ventriloquism. A lot of perception research has treated cue combination as a fixed rule: take a weighted average, done. Körding's work makes the rule conditional.

Weighted averages are what you do when, and only when, your brain has reason to believe there's a single thing out there throwing off multiple signals. That decision—about causal structure—drives the pattern of bias, the dependence on disparity, the unity judgments, and the quantitative goodness-of-fit. And it does so with a compact set of parameters that you can measure, interpret, and compare across observers.

There are always caveats. The model assumes Gaussian noise and a simple central prior; real environments can get messier. People's p common might change with context, learning, or instructions.

And the brain might approximate rather than compute the exact posterior. But in a tightly controlled setting, with individualized auditory cues and joint reports, a causal-inference observer nails what people do. It fuses when it should, separates when it must, and shades smoothly between those regimes as the evidence demands.

If you're curious where this goes next, imagine extending the same logic to time—are two flashes and one beep part of one event or two?—or to more than two senses, or to richer priors learned from natural scenes. The blueprint doesn't change: infer the structure, then combine accordingly. What Körding and colleagues showed is that once you make that first, simple decision explicit, a lot of perception that looked mysterious starts to make sense—and it comes with numbers you can test.

That's the sweet spot: story meeting substance, with a brain that binds the world together not by default, but by inference.

More in Psychology