Corpus linguistics for language teaching and learningA research agenda

Niall Curry, Tony McEneryView original
OverviewExpertonyx voice
Today I want to examine a very specific question raised explicitly in Curry and McEnery’s 2025 research agenda: what role should artificial intelligence play in advancing corpus-informed language teaching methodologies? Their answer, woven through the paper but crystallized in Research task 3, is both pragmatic and cautionary—artificial intelligence can lower long-standing barriers to data-driven learning, but only if pedagogy and ethics lead, and attested corpora remain the reference point. Let me ground us briefly. Over the last forty years, corpus linguistics has shifted language teaching from intuition to evidence, with direct applications like data-driven learning, abbreviated as DDL, and indirect applications spanning lexicography, reference grammars, materials, assessment, and teacher education. The initial revolution of the 1990s rode on accessible concordancers like WordSmith Tools and, later, AntConc; the more recent renewal, as the authors argue, comes from digital pedagogies and computational methods. But the research-practice gap persists. Meta-analyses by Boulton and Cobb and by Pérez-Paredes show good learning effects for concordancing and collocation work, yet adoption outside university contexts has lagged—training demands, time costs, and software friction have been stubborn obstacles. That’s the backdrop against which artificial intelligence arrives. So what exactly can artificial intelligence do here? The paper proposes we treat artificial intelligence as a potential mediator for data-driven learning rather than a replacement. One obvious affordance is usability: conversational interfaces can behave like a soft concordancer, allowing a teacher or learner to elicit pattern-like evidence with plain prompts. Lim and Wang show how generative artificial intelligence can be queried for distributional behavior of forms such as “interested in doing” versus “interested to do.” AntConc itself now includes integrated artificial intelligence functionality—Curry and McEnery point to this hybrid design as a natural place to study teacher and learner engagement. And there are nearer cousins already in use: ColloCaid offers feedforward collocation suggestions grounded in academic corpora, and Write and Improve delivers corpus-based corrective feedback at scale. Artificial intelligence can extend these ideas—surfacing candidate patterns, summarizing concordance lines, and scaffolding metalinguistic reflection—without sending every teacher to a keyword in context terminal on day one. But—and the authors say this plainly—artificial intelligence does not supply attested language in use. Generative systems synthesize; they do not sample. That distinction matters. Data-driven learning’s epistemic core is exposure to authentic distributions: frequency, collocational profiles, register variation. If we swap those for synthetic outputs, we’ve changed the method. Lin argues, and Curry and McEnery agree, that artificial intelligence is not a data-driven learning replacement. It’s a facilitator that still needs a corpus backbone. The practical implication is clear: any artificial intelligence-mediated workflow should route back to verifiable evidence from principled corpora. Let a model propose, but let Sketch Engine, the British National Corpus, or comparable resources decide. This is where pedagogy comes back to center stage. The paper is frank about data-driven learning’s lingering theoretical gap and insists that artificial intelligence integration be judged against explicit pedagogical criteria. They propose three tests: is the approach evidence-grounded, is the application innovative rather than gimmicky, and do learners actually build language competence and skills? That third point isn’t trivial. Digital pedagogy research the authors cite shows technology can boost engagement and autonomy—Croxton reports higher persistence, and Godwin-Jones links informal digital learning with self-direction—but access constraints and ethics routinely surface as counterweights. Sharkey’s cautionary work on chatbots—learners forming emotional bonds with artificial interlocutors—signals unanticipated socio-affective consequences in classrooms. So the research agenda asks for studies on artificial intelligence in data-driven learning that surface, not sidestep, these issues. What would a responsible methodology look like? Curry and McEnery sketch one. Start by situating artificial intelligence inside a corpus-centric tool that teachers already use—AntConc is their concrete example. Then investigate teacher and learner engagement through semi-structured interviews, with analysis via corpus-assisted grounded theory and top-down thematic coding. A dual lens matters here: the corpus view makes recurring beliefs and anxieties visible at scale, while grounded theory keeps the pedagogy in focus. The study’s lens should include metacognitive development as an outcome—Mizumoto explicitly ties data-driven learning and generative artificial intelligence to metacognitive resource use. If artificial intelligence can help learners plan searches, monitor hypothesis-testing across concordance lines, and evaluate paraphrases against distributions, then it’s not just “convenience tech”—it’s a scaffold for self-regulation. Of course, the benefits must be weighed against two classes of risk the authors emphasize: reliability and bias. Putland and colleagues show how artificial intelligence systems can produce confident but flawed analyses; Yuan and co-authors underscore instability and hallucination in academic-writing support. The Korean case that the paper discusses more broadly is also relevant: Choi documents how voice chatbots in South Korea reproduce native-speakerism. If artificial intelligence systems encode prescriptive ideologies or inner-circle norms by default, then uncritical classroom adoption collapses decades of progress on English as an international language. That’s not a marginal concern. The agenda explicitly ties artificial intelligence research to wider questions of representation raised in their research task 4 on pedagogical corpora: who gets represented in our data, and how? Let me get concrete about workflows—because this is where artificial intelligence can make a positive difference without eroding data-driven learning’s evidential spine. One, use artificial intelligence for query formulation. Teachers often struggle to articulate regular expression or wildcard queries; a model can translate a natural-language question—“show past passive forms of reporting verbs in research articles”—into a corpus query, then the human verifies results against concordances. Two, use artificial intelligence for concordance synthesis. Summarizing fifty keyword in context lines into salient collocational frames saves time while preserving attestedness. The model drafts the summary; the teacher cross-checks two or three anchor examples. Three, use artificial intelligence for contrastive prompts. In plurilingual or English for academic purposes settings, a model can propose cross-linguistic hypotheses—“do French abstracts prefer nominalization over English in move two?”—which learners then test in comparable or parallel corpora. The discovery remains empirical; artificial intelligence supplies the nudge. And four, embed metacognitive checkpoints. Ask learners to articulate, in one or two sentences, how the artificial intelligence’s suggestion matched—or failed to match—the corpus evidence. Those micro-reflections are the metacognitive resource-use moments Mizumoto highlights. Now, where does this leave indirect applications? The paper doesn’t position artificial intelligence as a panacea for lexicography, materials, or assessment, but the same principles apply. For materials, publishers already consult large corpora and market data; Curry, Love, and Goodman’s work with coursebook stakeholders shows the tension between national-variety findings and global English as an international language needs. If artificial intelligence is used to triage candidate examples or to align register profiling with Common European Framework of Reference targets—McCarthy’s learner-corpus work on Common European Framework of Reference calibration comes to mind—then human editorial judgement must still arbitrate inclusion, representation, and pedagogical fit. For assessment, learner corpora such as the Cambridge Learner Corpus or the Trinity Lancaster Corpus support empirical criteria development; artificial intelligence could speed annotation or pattern discovery, but test validity depends on transparent, auditable evidence. Xi’s question—what can corpus linguistics offer to assessment—precedes any promise of automation. The paper’s stance is consistent: artificial intelligence may accelerate workflows, yet the constructs and corpora define the measurement. A moment on context, because Curry and McEnery make it central in research task 2. They propose South Korea as a testbed for aligning corpus-informed methods with national curricula and affective aims. That has clear artificial intelligence implications. If high-stakes exams narrow communicative practice, artificial intelligence tutors can look like a release valve—low-stakes interaction, personalized drills, real-time feedback. But Choi’s chatbot findings warn that these systems can encode native-speakerist norms that sit uneasily with equity goals, and the curriculum’s mindfulness focus raises separate questions about screen-mediated practice. A corpus approach to national assessments—comparing listening and reading language across the replaced National English Ability Test and the newer College Scholastic Ability Test using metrics like type–token ratio and Common European Framework of Reference mapping—could identify input gaps. Only then does it make sense to ask what an artificial intelligence tutor should practice, and which corpora its suggestions must reference. Sequence matters. All of this circles back to the authors’ call for participatory research design—research task 5. If artificial intelligence is going to mediate data-driven learning at scale, its design needs co-ownership. Teachers, learners, publishers, assessment developers, and toolmakers should be in the same room—literally or metaphorically—articulating what counts as acceptable evidence, how digital literacies will be taught, and where the human-in-the-loop sits. Curry and Mark’s workshops with teachers on corpus-informed materials show there is appetite for that conversation, and that stakeholders don’t always share priorities. Grounded interviews, iterative guidelines, and action cycles from participatory action research give a process for reconciliation. Without that, artificial intelligence will simply amplify whichever stakeholder already has the loudest voice. Let me highlight a few numbers that set expectations. The agenda lays out five research tasks, and only one is narrowly about artificial intelligence. Data-driven learning interventions in the literature range from one to sixteen weeks; any artificial intelligence-mediated study should plan on a similar timescale to capture delayed uptake—Boulton and Cobb stress delayed post-tests for exactly this reason. And the authors’ proposed interview-based study orients to three evaluation criteria—evidence, innovation, and learning outcomes—so success isn’t measured in clicks or prompts, but in demonstrable linguistic development and metacognitive growth. Where are the edges of the map? Three limitations loom. First, data provenance. If model outputs are used as input for learning, we need empirical checks on their relevance and register appropriateness; the authors explicitly call for research on whether generative outputs are suitable as input. Second, bias and representation. Their task on pedagogical corpora pushes us to rethink sampling frames; artificial intelligence systems trained on skewed distributions will quietly reinscribe those skews in suggestions and feedback. Third, access. Hockly and Dudeney’s work on digital divides reminds us that artificial intelligence that assumes constant connectivity, high-end devices, or subscription fees will exacerbate inequities—precisely the opposite of the “lower the barrier” story many of us want to tell. So where does that leave us? With a balanced, tractable brief. Use artificial intelligence where it can reduce friction—query formulation, pattern summarization, reflection prompts—while keeping attested corpora as the evidential anchor. Design classroom studies that examine not only performance but also metacognitive resource use, with interviews analyzed through corpus-supported grounded theory to surface tacit teacher and learner models. Tie any artificial intelligence-mediated data-driven learning to explicit pedagogy rather than tool-first enthusiasm. And do it in conversation with the people who will live with the results—teachers, learners, and, yes, the publishers and assessment developers who shape materials and tests. If we follow that script, artificial intelligence can help data-driven learning finally break through persistent adoption barriers without hollowing out its epistemic core. It becomes an accelerator for the agenda Curry and McEnery set across their five tasks—supporting plurilingual, context-sensitive, and ethically grounded corpus applications—rather than a shortcut that bypasses the hard questions. And that’s the role worth claiming.

Today I want to examine a very specific question raised explicitly in Curry and McEnery’s 2025 research agenda: what role should artificial intelligence play in advancing corpus-informed language teaching methodologies? Their answer, woven through the paper but crystallized in Research task 3, is both pragmatic and cautionary—artificial intelligence can lower long-standing barriers to data-driven learning, but only if pedagogy and ethics lead, and attested corpora remain the reference point.

Let me ground us briefly. Over the last forty years, corpus linguistics has shifted language teaching from intuition to evidence, with direct applications like data-driven learning, abbreviated as DDL, and indirect applications spanning lexicography, reference grammars, materials, assessment, and teacher education. The initial revolution of the 1990s rode on accessible concordancers like WordSmith Tools and, later, AntConc; the more recent renewal, as the authors argue, comes from digital pedagogies and computational methods.

But the research-practice gap persists. Meta-analyses by Boulton and Cobb and by Pérez-Paredes show good learning effects for concordancing and collocation work, yet adoption outside university contexts has lagged—training demands, time costs, and software friction have been stubborn obstacles. That’s the backdrop against which artificial intelligence arrives.

So what exactly can artificial intelligence do here? The paper proposes we treat artificial intelligence as a potential mediator for data-driven learning rather than a replacement. One obvious affordance is usability: conversational interfaces can behave like a soft concordancer, allowing a teacher or learner to elicit pattern-like evidence with plain prompts.

Lim and Wang show how generative artificial intelligence can be queried for distributional behavior of forms such as “interested in doing” versus “interested to do.” AntConc itself now includes integrated artificial intelligence functionality—Curry and McEnery point to this hybrid design as a natural place to study teacher and learner engagement. And there are nearer cousins already in use: ColloCaid offers feedforward collocation suggestions grounded in academic corpora, and Write and Improve delivers corpus-based corrective feedback at scale. Artificial intelligence can extend these ideas—surfacing candidate patterns, summarizing concordance lines, and scaffolding metalinguistic reflection—without sending every teacher to a keyword in context terminal on day one.

But—and the authors say this plainly—artificial intelligence does not supply attested language in use. Generative systems synthesize; they do not sample. That distinction matters.

Data-driven learning’s epistemic core is exposure to authentic distributions: frequency, collocational profiles, register variation. If we swap those for synthetic outputs, we’ve changed the method. Lin argues, and Curry and McEnery agree, that artificial intelligence is not a data-driven learning replacement.

It’s a facilitator that still needs a corpus backbone. The practical implication is clear: any artificial intelligence-mediated workflow should route back to verifiable evidence from principled corpora. Let a model propose, but let Sketch Engine, the British National Corpus, or comparable resources decide.

This is where pedagogy comes back to center stage. The paper is frank about data-driven learning’s lingering theoretical gap and insists that artificial intelligence integration be judged against explicit pedagogical criteria. They propose three tests: is the approach evidence-grounded, is the application innovative rather than gimmicky, and do learners actually build language competence and skills?

That third point isn’t trivial. Digital pedagogy research the authors cite shows technology can boost engagement and autonomy—Croxton reports higher persistence, and Godwin-Jones links informal digital learning with self-direction—but access constraints and ethics routinely surface as counterweights. Sharkey’s cautionary work on chatbots—learners forming emotional bonds with artificial interlocutors—signals unanticipated socio-affective consequences in classrooms.

So the research agenda asks for studies on artificial intelligence in data-driven learning that surface, not sidestep, these issues.

What would a responsible methodology look like? Curry and McEnery sketch one. Start by situating artificial intelligence inside a corpus-centric tool that teachers already use—AntConc is their concrete example.

Then investigate teacher and learner engagement through semi-structured interviews, with analysis via corpus-assisted grounded theory and top-down thematic coding. A dual lens matters here: the corpus view makes recurring beliefs and anxieties visible at scale, while grounded theory keeps the pedagogy in focus. The study’s lens should include metacognitive development as an outcome—Mizumoto explicitly ties data-driven learning and generative artificial intelligence to metacognitive resource use.

If artificial intelligence can help learners plan searches, monitor hypothesis-testing across concordance lines, and evaluate paraphrases against distributions, then it’s not just “convenience tech”—it’s a scaffold for self-regulation.

Of course, the benefits must be weighed against two classes of risk the authors emphasize: reliability and bias. Putland and colleagues show how artificial intelligence systems can produce confident but flawed analyses; Yuan and co-authors underscore instability and hallucination in academic-writing support. The Korean case that the paper discusses more broadly is also relevant: Choi documents how voice chatbots in South Korea reproduce native-speakerism.

If artificial intelligence systems encode prescriptive ideologies or inner-circle norms by default, then uncritical classroom adoption collapses decades of progress on English as an international language. That’s not a marginal concern. The agenda explicitly ties artificial intelligence research to wider questions of representation raised in their research task 4 on pedagogical corpora: who gets represented in our data, and how?

Let me get concrete about workflows—because this is where artificial intelligence can make a positive difference without eroding data-driven learning’s evidential spine. One, use artificial intelligence for query formulation. Teachers often struggle to articulate regular expression or wildcard queries; a model can translate a natural-language question—“show past passive forms of reporting verbs in research articles”—into a corpus query, then the human verifies results against concordances.

Two, use artificial intelligence for concordance synthesis. Summarizing fifty keyword in context lines into salient collocational frames saves time while preserving attestedness. The model drafts the summary; the teacher cross-checks two or three anchor examples.

Three, use artificial intelligence for contrastive prompts. In plurilingual or English for academic purposes settings, a model can propose cross-linguistic hypotheses—“do French abstracts prefer nominalization over English in move two?”—which learners then test in comparable or parallel corpora. The discovery remains empirical; artificial intelligence supplies the nudge.

And four, embed metacognitive checkpoints. Ask learners to articulate, in one or two sentences, how the artificial intelligence’s suggestion matched—or failed to match—the corpus evidence. Those micro-reflections are the metacognitive resource-use moments Mizumoto highlights.

Now, where does this leave indirect applications? The paper doesn’t position artificial intelligence as a panacea for lexicography, materials, or assessment, but the same principles apply. For materials, publishers already consult large corpora and market data; Curry, Love, and Goodman’s work with coursebook stakeholders shows the tension between national-variety findings and global English as an international language needs.

If artificial intelligence is used to triage candidate examples or to align register profiling with Common European Framework of Reference targets—McCarthy’s learner-corpus work on Common European Framework of Reference calibration comes to mind—then human editorial judgement must still arbitrate inclusion, representation, and pedagogical fit. For assessment, learner corpora such as the Cambridge Learner Corpus or the Trinity Lancaster Corpus support empirical criteria development; artificial intelligence could speed annotation or pattern discovery, but test validity depends on transparent, auditable evidence. Xi’s question—what can corpus linguistics offer to assessment—precedes any promise of automation.

The paper’s stance is consistent: artificial intelligence may accelerate workflows, yet the constructs and corpora define the measurement.

A moment on context, because Curry and McEnery make it central in research task 2. They propose South Korea as a testbed for aligning corpus-informed methods with national curricula and affective aims. That has clear artificial intelligence implications.

If high-stakes exams narrow communicative practice, artificial intelligence tutors can look like a release valve—low-stakes interaction, personalized drills, real-time feedback. But Choi’s chatbot findings warn that these systems can encode native-speakerist norms that sit uneasily with equity goals, and the curriculum’s mindfulness focus raises separate questions about screen-mediated practice. A corpus approach to national assessments—comparing listening and reading language across the replaced National English Ability Test and the newer College Scholastic Ability Test using metrics like type–token ratio and Common European Framework of Reference mapping—could identify input gaps.

Only then does it make sense to ask what an artificial intelligence tutor should practice, and which corpora its suggestions must reference. Sequence matters.

All of this circles back to the authors’ call for participatory research design—research task 5. If artificial intelligence is going to mediate data-driven learning at scale, its design needs co-ownership. Teachers, learners, publishers, assessment developers, and toolmakers should be in the same room—literally or metaphorically—articulating what counts as acceptable evidence, how digital literacies will be taught, and where the human-in-the-loop sits.

Curry and Mark’s workshops with teachers on corpus-informed materials show there is appetite for that conversation, and that stakeholders don’t always share priorities. Grounded interviews, iterative guidelines, and action cycles from participatory action research give a process for reconciliation. Without that, artificial intelligence will simply amplify whichever stakeholder already has the loudest voice.

Let me highlight a few numbers that set expectations. The agenda lays out five research tasks, and only one is narrowly about artificial intelligence. Data-driven learning interventions in the literature range from one to sixteen weeks; any artificial intelligence-mediated study should plan on a similar timescale to capture delayed uptake—Boulton and Cobb stress delayed post-tests for exactly this reason.

And the authors’ proposed interview-based study orients to three evaluation criteria—evidence, innovation, and learning outcomes—so success isn’t measured in clicks or prompts, but in demonstrable linguistic development and metacognitive growth.

Where are the edges of the map? Three limitations loom. First, data provenance.

If model outputs are used as input for learning, we need empirical checks on their relevance and register appropriateness; the authors explicitly call for research on whether generative outputs are suitable as input. Second, bias and representation. Their task on pedagogical corpora pushes us to rethink sampling frames; artificial intelligence systems trained on skewed distributions will quietly reinscribe those skews in suggestions and feedback.

Third, access. Hockly and Dudeney’s work on digital divides reminds us that artificial intelligence that assumes constant connectivity, high-end devices, or subscription fees will exacerbate inequities—precisely the opposite of the “lower the barrier” story many of us want to tell.

So where does that leave us? With a balanced, tractable brief. Use artificial intelligence where it can reduce friction—query formulation, pattern summarization, reflection prompts—while keeping attested corpora as the evidential anchor.

Design classroom studies that examine not only performance but also metacognitive resource use, with interviews analyzed through corpus-supported grounded theory to surface tacit teacher and learner models. Tie any artificial intelligence-mediated data-driven learning to explicit pedagogy rather than tool-first enthusiasm. And do it in conversation with the people who will live with the results—teachers, learners, and, yes, the publishers and assessment developers who shape materials and tests.

If we follow that script, artificial intelligence can help data-driven learning finally break through persistent adoption barriers without hollowing out its epistemic core. It becomes an accelerator for the agenda Curry and McEnery set across their five tasks—supporting plurilingual, context-sensitive, and ethically grounded corpus applications—rather than a shortcut that bypasses the hard questions. And that’s the role worth claiming.