Individual Differences in Inhibitory Control, Not Non-Verbal Number Acuity, Correlate with Mathematics Achievement
Here is the received wisdom: children who are better at judging which of two clouds of dots is more numerous are better at math. That finding has been replicated dozens of times. It pointed toward something exciting — an ancient, pre-verbal sense of quantity that might be the hidden foundation of mathematical ability. Researchers built assessments around it. They built training programs around it. Then Gilmore and colleagues looked more carefully at the task itself, and the effect fell apart. The idea behind the dot comparison task is straightforward. You show someone two arrays of dots — one red, one blue — and ask which has more. Do this across many trials at different ratios, and you get a measure of how precisely a person can discriminate numerical quantities without using words or symbols. This precision is captured by a parameter called w, which tracks the fuzziness of numerical representations. The lower the w, the sharper the number sense. The correlation between w and accuracy on the task was nearly perfect — r equals negative 0.94 — so the two measures essentially stand in for each other. What fascinated researchers was that this capacity, the Approximate Number System, or ANS, seemed to predict formal mathematics achievement. The implication was tantalizing: maybe the raw, pre-linguistic sense of number is what formal math is built on. If so, you could potentially train it.
But designing that task turns out to be harder than it looks. When you put two dot clouds side by side, the more numerous one tends to take up more space, contain larger dots, and cover more total area. So a participant who isn't counting at all — who's just reacting to overall visual magnitude — will still get a lot of trials right. To control for this, researchers split trials into two types. Congruent trials are ones where the visual cues line up with numerosity: the bigger cloud looks bigger. Incongruent trials flip it: the more numerous cloud has smaller dots and takes up less space, so the visual signal actually points you in the wrong direction. To get an incongruent trial right, you have to suppress your instinct to grab the visually larger array and respond on the basis of number alone. Gilmore and colleagues recognized this as inhibitory control — the same cognitive demand as a Stroop task, where you have to name the ink color while ignoring what the word says. The behavioral fingerprints of that suppression are exactly what you'd expect. In Experiment 1, which tested eighty children between ages four and twelve, mean accuracy on congruent trials was 0.82. On incongruent trials, it dropped to 0.51 — barely above chance.
Correct responses on incongruent trials took an average of 1,413 milliseconds, compared to 1,067 milliseconds on congruent trials. And here's the telling detail: on incongruent trials, correct responses were actually slower than incorrect ones. Children who got those trials wrong had latched onto the fast, automatic visual response. Children who got them right had taken the time to override it. That is the signature of inhibition at work, not number sense. Now here is where the mathematics connection enters. When Gilmore and colleagues correlated dot task performance with standardized arithmetic scores and split by trial type, the picture became unambiguous. Overall dot-comparison accuracy correlated with math achievement at r equals 0.57. But when they pulled the two trial types apart, the correlation for incongruent trials alone was 0.55, and the correlation for congruent trials alone was essentially zero — r equals 0.03, with a p-value of 0.80. A Williams-Steiger test confirmed those two correlations were significantly different from each other. The cleaner measure of ANS acuity — performance on trials where visual cues and numerosity agree — predicted nothing about mathematics achievement. Only the inhibition-loaded trials did.
Experiment 2 tested seventy-one children between ages seven and ten and added a direct measure of inhibitory control: the NEPSY-II Inhibition subtest, a standardized task that asks children to switch between naming shapes and applying a rule about them — a classic executive function measure. Mathematics achievement was assessed with the WIAT-II UK Numerical Operations subtest. The researchers ran hierarchical regression analyses entering predictors in different orders. When dot-comparison ANS acuity was entered first, it significantly predicted math scores — a beta of negative 0.35, accounting for twelve percent of the variance. But when NEPSY-II inhibition and naming scores were added to the model, inhibition emerged as the significant predictor, and the ANS acuity coefficient shrank to negative 0.16 and lost significance. Flipping the order of entry produced the mirror image: inhibition entered first was a strong predictor, accounting for twenty-six percent of variance on its own, and adding dot-comparison w afterward improved the model by only two percent — a non-significant gain. The naming score, which controlled for the speed demands of the NEPSY task, did not drive the effect. It was specifically inhibitory control.
Gilmore and colleagues also reanalyzed a prior dataset from Inglis and colleagues and found the same pattern holding in that independent sample: incongruent trial performance correlated with math achievement at r equals 0.34, while congruent trial performance produced a non-significant r of negative 0.22. What this means is that the widely reported link between dot comparison performance and mathematics achievement was measuring something real — just not what researchers thought. The children who scored higher on the dot task weren't the ones with the sharpest number sense. They were the ones who could override a misleading visual signal. That is a domain-general cognitive skill. It has nothing specifically to do with quantity representation. This matters enormously for what followed from the original finding. Websites offering ANS training were built on the premise that sharpening the ANS would improve mathematics learning. Assessments were designed around dot comparison tasks as measures of foundational numerical ability. If those tasks are primarily tapping inhibitory control, then training performance on them may have been training inhibitory control all along, or training neither, while the underlying ANS went unmeasured and unimproved.
The harder question the paper opens, but doesn't fully close, is what the right conclusion is for education. If inhibitory control is what's actually correlated with mathematics achievement, then maybe that's the lever worth pulling. Domain-general executive functions — the ability to hold information in mind, to switch between rules, to suppress automatic responses — have been implicated in academic learning broadly. A child who struggles to inhibit a salient visual distractor may also struggle to hold a borrowing procedure in working memory, or to suppress an overlearned arithmetic fact when the problem demands a different approach. These are speculative connections, and Gilmore and colleagues are careful not to overclaim them. What they do claim, precisely and with evidence, is that the dot comparison task as it has been standardly used cannot tell you what you want to know about the ANS. Any correlation it produces with math achievement is tangled up with inhibitory demands that have no necessary connection to number representations at all.
The call to action from the paper is pointed: researchers need task designs that actually separate ANS acuity from inhibitory control — formats where continuous visual variables are controlled without creating conflicting signals that require suppression. Until that separation is achieved, every study that uses a congruent-incongruent dot comparison design and finds a correlation with math scores is catching both things in the same net and reporting it as one fish. Gilmore and colleagues untied that knot. What researchers and educators do next with the untangled threads is the real test of whether the finding sticks. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.
Related lectures
- Real-time tentative assessment of the epidemiological characteristics of novel coronavirus infections in Wuhan, China, as at 22 January 2020
- The impact of non-pharmaceutical interventions on SARS-CoV-2 transmission across 130 countries and territories
- Community Transmission of Severe Acute Respiratory Syndrome Coronavirus 2, Shenzhen, China, 2020
- High-Resolution Measurements of Face-to-Face Contact Patterns in a Primary School
- Early dynamics of transmission and control of COVID-19: a mathematical modelling study
- Serial Interval of COVID-19 among Publicly Reported Confirmed Cases