Causal reasoning with forces
Think about how you learn new causal facts in the world. You sand a piece of wood, sanding makes dust, dust makes you sneeze — so sanding makes you sneeze. Sometimes the links don’t match, though: a fertilizer helps plants grow, and extra growth prevents erosion, so the fertilizer prevents erosion.
That step, where you combine different kinds of links, is called relation composition. Instead of simple transitivity, you are composing "cause," "allow," and "prevent" into a new, single relation between the first thing and the last.
There’s a long-running debate about how the mind does this. One camp, the mental model tradition, says we translate language into abstract structures that encode what must be true, then operate over those. Another, the causal model tradition, leans on probabilistic dependencies among variables.
Wolff and colleagues offer a third path: the force theory. It says our causal concepts rely on iconic, perceptual codes — essentially little vector diagrams — that a physics-like simulator can run. In this view, we can compose causal relations because the codes behind "cause," "allow," and "prevent" are like force configurations that can be added, subtracted, and played forward.
Here’s the core representation. Every simple causal link has an affector A pushing on a patient P toward or away from an endstate E — think of E as a target location, not just a direction. The patient has its own tendency, described as a vector P.
The affector contributes a vector A. If you add up the relevant forces, you get a resultant R. If R points toward E, the patient reaches the goal; if not, it doesn’t.
With that in place, the three main causal meanings emerge from three cues people intuitively use: Is the patient already tending toward the goal? Do the affector and patient line up or conflict? And does the endstate get reached?
In a CAUSE scene, the patient wasn't headed there, the affector pushes against the patient’s tendency, and the endstate is reached. In ALLOW or ENABLE, the patient was already going that way, the affector helps, and the endstate is reached. In PREVENT, the patient was heading toward the goal, the affector opposes it, and the endstate is not reached.
People are often vague about exact magnitudes, and the theory accommodates that; directions and rough strength matter most.
Composition is just vector bookkeeping. To summarize a chain, take the affector from the first link, the endstate from the last, and sum the intervening patient forces. That holds true whether forces are transmitted — like A causes B, and B causes C — or removed — like A prevents B, and B prevents C.
The interesting wrinkle is double prevention. If A prevents B and B prevents C, the net effect can be that A allows C, but in some scenarios, A ends up causing C. If you run the physics many times with different plausible magnitudes, you see a stable split: ALLOW about 62 percent of the time, CAUSE about 38 percent.
Wolff's team got the same split in two ways, once by analytic integration over the space of magnitudes, and once by a stochastic simulation that samples many combinations and tallies outcomes. Some compositions are tighter. Cause composed with Prevent tends to yield Prevent essentially all the time.
Prevent composed with Cause yields Prevent only about 37 percent of the time, with the rest falling into indeterminate territory when the vector sizes leave the conclusion underspecified.
So, does anyone actually reason this way? Wolff and colleagues built a perceptual testbed to find out. They instantiated those force-theory configurations in a real physics engine and rendered them as short animations of colored cars bumping and blocking each other.
ALLOW was constructed as two nested PREVENT relations, just as the analysis suggests. The motion was handled by Havok's dynamics; the point was to make the forces concrete and controllable. Emory undergraduates — 24 in a first pass to check the basic scenes, 25 in the main test — watched pairs of clips that together formed a chain like CAUSE and CAUSE or PREVENT and ALLOW.
After each, they picked the best description of the relation between the first and last cars: caused, allowed, prevented, or none of the above.
Across all nine chain types, the force theory's predicted compositions matched the modal choices that people made. Take a couple of examples. When both links were ALLOW, participants chose the sentence with allow; the chi-square statistic for that preference was strong, with a chi-square value of 21.56 and a p-value below 0.0001.
When both links were CAUSE, they picked cause, with a chi-square value of 24.56 and again with a p-value below 0.0001. And when the chain was ALLOW then PREVENT, they chose prevent, with a chi-square value of 13.88 and a p-value under 0.01. The one designed ambiguity, PREVENT then PREVENT — the double prevention — produced a split between cause and allow with no single winner, just as the force theory suggests it should.
The big picture isn’t just that the right labels were chosen. It’s that animated scenes, built from vector-like parts and run in a physics engine, led observers to the same composed conclusions the vector math predicts. That’s a rare kind of perceptual grounding.
Those probabilistic double-prevention outcomes weren’t an afterthought; they’re built into the theory. The team’s calculus-based and iterative simulations each landed on the same sixty-two to thirty-eight split between ALLOW and CAUSE, and participants’ willingness to give either answer in the ambiguous chain lined up with that indeterminacy. That’s a neat payoff: the same representation that explains crisp cases also explains why some compositions aren’t crisp.
Next, the authors left the toy world of cars and ramps and asked whether the same machinery shows up when people read abstract causal statements. Here, they pitted the force theory against the causal model and mental model accounts, doing it in three increasingly demanding tests, always keeping an eye on two ways to score predictions: the dominant, modal conclusion and the full distribution of responses. They also took seriously a known bias — Sloman and colleagues’ "matching" or atmosphere effect — where people tend to use the same polarity words in their conclusions that they saw in the premises.
That bias can pull people toward negations or away from them, sometimes helping a theory’s fit and sometimes hurting it.
In the second experiment, they borrowed materials from Goldvarg and Johnson-Laird: two-premise chains phrased in abstract language, including negations like "No learning causes anxiety." The data had a clear structure. People said CAUSE eighty-five percent of the time for a cause–cause chain. They said PREVENT one hundred percent of the time for cause–prevent.
For allow–allow, ninety-five percent chose allow. Some cases were mixed, like allow–cause, where sixty-five percent went with allow and thirty-five percent with cause. Negated premises could shift the balance: for allow–not-cause, prevent was the top answer at seventy percent, even though the force theory, taken strictly, leans toward allow for that configuration.
When you step back, however, the scoring is straightforward: both the force theory and the causal model account correctly predicted the dominant answer for fourteen of sixteen compositions; the mental model account got thirteen of sixteen. The distribution-sensitive analyses told a similar story. The headline isn’t that one theory trounced the others.
It’s that the force theory, built from perceptual pieces, held its own with a classic probabilistic account on purely linguistic tasks.
Then came real-world content. In the third experiment, the premises were drawn from web text — "Economic freedom causes wealth," that sort of thing — and the space of chain types doubled. Forty participants saw six exemplars of each of thirty-two compositions; all plausible conclusions were on the table, including explicit negations like not-cause.
With messier content and six times more trials, the force and causal model accounts actually pulled further ahead. Each correctly predicted thirty of thirty-two dominant conclusions; the mental model account got twenty-six of thirty-two. A Friedman test comparing the counts across theories found that the top two outperformed the mental model account, with a chi-square value of eight and a p-value of 0.018.
A couple of chains with double negations — essentially "not-cause" combined with allow or prevent — tripped up all three theories, and the authors traced those misses to the matching bias built into the task. When they allowed paraphrases that softened the negation wording, the force theory, and often the causal model account, could anticipate the observed choices. One replication from the animations popped out here too: in prevent–prevent chains, the most common conclusion was allow, not prevent, echoing the double-prevention logic observed in the lab.
Finally, the fourth experiment stretched the chains to three premises. Twenty-four participants worked through real-world triples like allow–cause–cause and cause–not-cause–cause. The key question was whether a vector-summation scheme scales.
It did. The force and causal model theories each nailed twenty-three of twenty-five dominant conclusions; the mental model theory got fourteen of twenty-five. The statistical comparison favored the first two again, with a chi-square value of eighteen and a p-value under 0.001.
The same tricky double-negation chains remained problematic for all three theories. One chain, cause–not-cause–prevent, showed the force theory predicting a split — about sixty percent cause and forty percent allow — while participants leaned more toward allow, around forty-two percent, with cause much lower. As before, relaxing the exact phrasing to reduce atmosphere effects brought the predictions and observations closer.
Underneath the numbers is a simple, appealing idea: abstract language can recruit the same force-like representations we use to parse physical scenes. That redeployment hypothesis fits the animation results, the abstract chains, and the real-world premises. It also fits our intuitions about cross-type composition.
If A causes B and B prevents C, you don’t need a probability table to conclude that "A prevents C." You can picture the forces and add them up. If A prevents B and B prevents C, you can picture why that often feels like "A allows C," with room for "A causes C" when the magnitudes shift. Even negation gets a concrete gloss. "A causes not-B" is just "A prevents B." And a "cause by omission" — not-A causes C — often translates to a double prevention, where the absence of a stopper removes a block.
There are limits. The force theory can overreach or undershoot in some double-prevention contexts, especially when wording invites matching biases. Some chains stumped everyone, hinting at places where world knowledge, non-iconic abstractions, or pragmatic choices steer people’s answers.
But across concrete animations and abstract text, the through-line holds: a compact, vector-based code predicts how people compose causal relations, and it does so as well as, and sometimes better than, more traditional accounts.
If you want the broader moral, it’s this. Human causal cognition seems to run on two engines at once. One is the formal, algebraic side — probabilities, dependencies, and the logic of interventions.
The other is sensorimotor — forces, tendencies, and endstates you can picture. Wolff’s work shows the second engine doesn’t stop at perception. It scales up, it composes, and it can carry abstract language on its back. That makes the space between seeing and thinking a lot smaller than it looks.
Related lectures
- Insomnia and the risk of depression: a meta-analysis of prospective cohort studies
- Mirror-Induced Behavior in the Magpie (Pica pica): Evidence of Self-Recognition
- The cross-national epidemiology of social anxiety disorder: Data from the World Mental Health Survey Initiative
- The Natural Statistics of Audiovisual Speech
- The Small World of Psychopathology
- Health-related quality of life in parents of school-age children with Asperger syndrome or high-functioning autism