An ERP study on L2 syntax processingWhen do learners fail?

Nienke Meulman, Laurie A. Stowe, Simone Sprenger, Moniek Bresser, Monika S. SchmidView original
OverviewBalancedmarcus voice
Here is a group of people who live in Dutch. They work in Dutch, argue in Dutch, and probably count in Dutch. They have been immersed for years, and if you hand them a grammar test, they score near the top. Ask them whether a noun takes a common or neuter determiner, and many of them get it right. By every standard measure, they appear nearly native. But put them in an EEG cap and show them sentences with gender errors, and something goes quiet. The brain signal that should fire — the one that fires reliably in every native speaker — simply doesn't appear. Knowing a rule and processing it in real time, it turns out, are completely different things. And Meulman and colleagues made that gap visible, millisecond by millisecond. The tool they used is event-related brain potentials, or ERPs. These are voltage fluctuations recorded from the scalp that track the brain's response to incoming language in real time. The component they focused on is called the P600, which is a positive-going voltage wave that peaks roughly six hundred milliseconds after encountering a grammatical violation. Think of it as the brain's grammar flag. When a native speaker hits a syntactic error, the P600 fires. Meulman and colleagues used it as a neurophysiological marker of native-like syntactic sensitivity, measuring amplitudes in a six hundred to twelve hundred millisecond window after the critical word. The question was whether highly proficient, immersed Romance language learners of Dutch would show the same flag. To understand why Dutch is a particularly hard test case, you have to understand what makes its gender system unusual. Dutch assigns every noun to one of two grammatical classes: common or neuter. This assignment determines which determiner you use and how adjectives inflect. So far, this is not so different from French or Spanish. The brutal part is that Dutch gender assignment is largely arbitrary at the lexical level. Romance languages give you morphophonological cues — patterns in word endings that predict gender with reasonable reliability. Dutch, according to Meulman and colleagues, is "generally regarded as having an opaque gender system." Some cues exist, but they cover only a small slice of the vocabulary. For the rest, you have to tag each noun individually and remember it. There's no shortcut. This matters for learners who come from Romance language backgrounds, because having gender in your first language doesn't automatically help you. What transfers isn't the abstract concept of grammatical gender; it's the processing routines tied to how gender works in your native language. When those routines don't map cleanly onto Dutch, they can interfere rather than assist. Meulman and colleagues note that Sabourin and Stowe made exactly this argument. L1 routines that are similar to L2 routines can help, but when they diverge, they can get in the way. To isolate the problem, the study included a second grammatical structure as an internal control: non-finite verb agreement violations. This is a more regular, transparent construction in Dutch — a cleaner rule with more consistent cues. If learners failed at both, that would suggest a general late-learner limitation. If they failed only at gender, that would point specifically to the combination of first and second language mismatch and target-language opacity. Nineteen Romance learners of Dutch and nineteen native Dutch speakers completed the experiment, reading or listening to sentences with either gender violations or verb violations while their EEG was recorded across fifty-four scalp electrodes. The results showed a clean dissociation. For non-finite verb violations, native speakers displayed the full expected pattern: an early negativity in the three hundred to five hundred millisecond window followed by a robust P600. The learners also showed a reliable P600 to verb violations — their brains flagged those errors at the late reanalysis stage, statistically matching the native pattern. The F-statistic for learners in the six hundred to twelve hundred millisecond window reached fourteen point sixteen on one degree of freedom against eighteen, with a p-value below 0.001. For gender violations, the picture was entirely different. Native speakers produced a clear P600 in temporal, parietal, and occipital regions — the parietal F alone reached thirty-eight point twenty, with a p-value below 0.001. Learners produced nothing. No reliable correctness effect appeared anywhere in the six hundred to twelve hundred millisecond window, with all follow-up F-values below one point eighty-one. At the individual level, the paper puts it plainly: basically none of the learners showed any sensitivity to gender violations in the ERP. This wasn't averaging over opposite effects. It was absence. Then comes the modality question — and it's a clever one. Maybe the issue wasn't the grammar itself, but the difficulty of the task. The standard visual presentation used in ERP studies, called rapid serial visual presentation, or RSVP, flashes words quickly one at a time. It's artificial. What if using natural auditory speech — the modality in which these learners actually live their daily Dutch lives — reduced task demands enough to bring out latent native-like processing? Meulman and colleagues ran both modalities within the same participants, allowing them to act as their own controls. The answer was: partially, and only for verbs. For non-finite verb violations, learners showed an early N400-like negativity — that earlier negative wave associated with processing difficulty — but only in the auditory condition. The correctness effect in the three hundred to five hundred millisecond window reached significance for auditory stimuli, with an F of six point eighteen and a p-value of 0.046, but not for visual stimuli, where the F was zero point forty-three. So, the more naturalistic listening condition did reveal some early sensitivity in learners for the transparent construction. For gender, modality changed nothing. Learners failed to show a reliable P600 to gender violations in both written and auditory conditions. The conclusion is hard to escape: the gender problem is not about task difficulty. Making the presentation easier and more natural didn't rescue online gender processing. Something deeper is at work. The individual differences analysis makes the point even sharper. Meulman and colleagues tested every obvious candidate: age of acquisition, length of residence, proficiency on a C-test, and offline gender knowledge — to see which predicted native-like P600 responses. None of them did for gender. Age of acquisition, the factor that dominates so much second-language acquisition research, had no significant effect. Neither did years of living in the Netherlands. Neither did proficiency scores. Neither did how well learners performed on an explicit gender knowledge test. The one variable that mattered was the amount of daily language use — and only for verb violations. A regression model explained thirty-three point seven percent of variance in P600 magnitude, with daily use of Dutch positively predicting P600 size for the verb condition, with an R-squared of zero point thirty-two and a p-value of 0.011. But for gender, daily language use predicted essentially nothing, with an R-squared of zero point zero one and a p-value of 0.756. The interaction term in the regression — the coefficient for gender structure multiplied by daily language use — was strongly negative, meaning that whatever benefit daily language use provided for verb processing was reversed for gender. Length of residence did correlate with offline gender knowledge, with a correlation of zero point fifty-five, and a p-value of 0.014 — so more time in the country did improve what learners consciously knew about Dutch gender. It just had no relationship to the P600, with a correlation of negative zero point eleven. You can spend years building explicit knowledge of Dutch gender, score better and better on tests, and your brain's real-time response to a gender error stays flat. This is the dissociation the paper exists to document. Offline knowledge and online processing come apart. Learners who pass a gender test, who correctly assign determiners when asked, and who have demonstrably learned the rule — those same learners show no neural signature of catching gender errors as sentences unfold. Standard fluency measures, the ones teachers and researchers rely on, can mask a fundamentally different processing system underneath. What Meulman and colleagues show is that for grammatical structures that are opaque and arbitrarily marked — where there's no reliable formal cue in the language, and where the learner's native processing routines don't transfer cleanly — late second-language learners may never fully automate the kind of rapid, implicit error detection that native speakers perform without effort. More immersion helps for some things. More daily use helps for some things. But Dutch gender, it seems, remains stubbornly outside the reach of all of those factors. The brain knows what it doesn't know, even when the person thinks they've learned it. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

Here is a group of people who live in Dutch. They work in Dutch, argue in Dutch, and probably count in Dutch. They have been immersed for years, and if you hand them a grammar test, they score near the top. Ask them whether a noun takes a common or neuter determiner, and many of them get it right. By every standard measure, they appear nearly native. But put them in an EEG cap and show them sentences with gender errors, and something goes quiet. The brain signal that should fire — the one that fires reliably in every native speaker — simply doesn't appear. Knowing a rule and processing it in real time, it turns out, are completely different things. And Meulman and colleagues made that gap visible, millisecond by millisecond. The tool they used is event-related brain potentials, or ERPs. These are voltage fluctuations recorded from the scalp that track the brain's response to incoming language in real time. The component they focused on is called the P600, which is a positive-going voltage wave that peaks roughly six hundred milliseconds after encountering a grammatical violation. Think of it as the brain's grammar flag. When a native speaker hits a syntactic error, the P600 fires. Meulman and colleagues used it as a neurophysiological marker of native-like syntactic sensitivity, measuring amplitudes in a six hundred to twelve hundred millisecond window after the critical word.

The question was whether highly proficient, immersed Romance language learners of Dutch would show the same flag. To understand why Dutch is a particularly hard test case, you have to understand what makes its gender system unusual. Dutch assigns every noun to one of two grammatical classes: common or neuter. This assignment determines which determiner you use and how adjectives inflect. So far, this is not so different from French or Spanish. The brutal part is that Dutch gender assignment is largely arbitrary at the lexical level. Romance languages give you morphophonological cues — patterns in word endings that predict gender with reasonable reliability. Dutch, according to Meulman and colleagues, is "generally regarded as having an opaque gender system." Some cues exist, but they cover only a small slice of the vocabulary. For the rest, you have to tag each noun individually and remember it. There's no shortcut. This matters for learners who come from Romance language backgrounds, because having gender in your first language doesn't automatically help you. What transfers isn't the abstract concept of grammatical gender; it's the processing routines tied to how gender works in your native language. When those routines don't map cleanly onto Dutch, they can interfere rather than assist. Meulman and colleagues note that Sabourin and Stowe made exactly this argument. L1 routines that are similar to L2 routines can help, but when they diverge, they can get in the way.

To isolate the problem, the study included a second grammatical structure as an internal control: non-finite verb agreement violations. This is a more regular, transparent construction in Dutch — a cleaner rule with more consistent cues. If learners failed at both, that would suggest a general late-learner limitation. If they failed only at gender, that would point specifically to the combination of first and second language mismatch and target-language opacity. Nineteen Romance learners of Dutch and nineteen native Dutch speakers completed the experiment, reading or listening to sentences with either gender violations or verb violations while their EEG was recorded across fifty-four scalp electrodes. The results showed a clean dissociation. For non-finite verb violations, native speakers displayed the full expected pattern: an early negativity in the three hundred to five hundred millisecond window followed by a robust P600. The learners also showed a reliable P600 to verb violations — their brains flagged those errors at the late reanalysis stage, statistically matching the native pattern. The F-statistic for learners in the six hundred to twelve hundred millisecond window reached fourteen point sixteen on one degree of freedom against eighteen, with a p-value below 0.001.

For gender violations, the picture was entirely different. Native speakers produced a clear P600 in temporal, parietal, and occipital regions — the parietal F alone reached thirty-eight point twenty, with a p-value below 0.001. Learners produced nothing. No reliable correctness effect appeared anywhere in the six hundred to twelve hundred millisecond window, with all follow-up F-values below one point eighty-one. At the individual level, the paper puts it plainly: basically none of the learners showed any sensitivity to gender violations in the ERP. This wasn't averaging over opposite effects. It was absence. Then comes the modality question — and it's a clever one. Maybe the issue wasn't the grammar itself, but the difficulty of the task. The standard visual presentation used in ERP studies, called rapid serial visual presentation, or RSVP, flashes words quickly one at a time. It's artificial. What if using natural auditory speech — the modality in which these learners actually live their daily Dutch lives — reduced task demands enough to bring out latent native-like processing? Meulman and colleagues ran both modalities within the same participants, allowing them to act as their own controls.

The answer was: partially, and only for verbs. For non-finite verb violations, learners showed an early N400-like negativity — that earlier negative wave associated with processing difficulty — but only in the auditory condition. The correctness effect in the three hundred to five hundred millisecond window reached significance for auditory stimuli, with an F of six point eighteen and a p-value of 0.046, but not for visual stimuli, where the F was zero point forty-three. So, the more naturalistic listening condition did reveal some early sensitivity in learners for the transparent construction. For gender, modality changed nothing. Learners failed to show a reliable P600 to gender violations in both written and auditory conditions. The conclusion is hard to escape: the gender problem is not about task difficulty. Making the presentation easier and more natural didn't rescue online gender processing. Something deeper is at work. The individual differences analysis makes the point even sharper. Meulman and colleagues tested every obvious candidate: age of acquisition, length of residence, proficiency on a C-test, and offline gender knowledge — to see which predicted native-like P600 responses. None of them did for gender.

Age of acquisition, the factor that dominates so much second-language acquisition research, had no significant effect. Neither did years of living in the Netherlands. Neither did proficiency scores. Neither did how well learners performed on an explicit gender knowledge test. The one variable that mattered was the amount of daily language use — and only for verb violations. A regression model explained thirty-three point seven percent of variance in P600 magnitude, with daily use of Dutch positively predicting P600 size for the verb condition, with an R-squared of zero point thirty-two and a p-value of 0.011. But for gender, daily language use predicted essentially nothing, with an R-squared of zero point zero one and a p-value of 0.756. The interaction term in the regression — the coefficient for gender structure multiplied by daily language use — was strongly negative, meaning that whatever benefit daily language use provided for verb processing was reversed for gender. Length of residence did correlate with offline gender knowledge, with a correlation of zero point fifty-five, and a p-value of 0.014 — so more time in the country did improve what learners consciously knew about Dutch gender. It just had no relationship to the P600, with a correlation of negative zero point eleven. You can spend years building explicit knowledge of Dutch gender, score better and better on tests, and your brain's real-time response to a gender error stays flat.

This is the dissociation the paper exists to document. Offline knowledge and online processing come apart. Learners who pass a gender test, who correctly assign determiners when asked, and who have demonstrably learned the rule — those same learners show no neural signature of catching gender errors as sentences unfold. Standard fluency measures, the ones teachers and researchers rely on, can mask a fundamentally different processing system underneath. What Meulman and colleagues show is that for grammatical structures that are opaque and arbitrarily marked — where there's no reliable formal cue in the language, and where the learner's native processing routines don't transfer cleanly — late second-language learners may never fully automate the kind of rapid, implicit error detection that native speakers perform without effort. More immersion helps for some things. More daily use helps for some things. But Dutch gender, it seems, remains stubbornly outside the reach of all of those factors. The brain knows what it doesn't know, even when the person thinks they've learned it. This lecture was created by ennepō. Go to https://ennepo.ai to Discover, Create and Follow the latest research in your field. Read when you can. Listen when you want to.

More in Neuroscience