Reproducible research practices and transparency across linguistics

Agata Bochyńska, Liam Keeble, Caitlin Halfacre, Joseph V. Casillas, Irys-Amélie Champagne et al. (+4)View original
HighlightsAccessiblealloy voice
Today we’re looking at a health check on how open and reproducible linguistics research really is. The main finding is stark: across a random sample of 600 linguistics journal articles spanning two eras, only about a third were openly accessible, fewer than one in ten shared their materials, data, or analysis scripts, no articles were preregistered, only 1% were replications, and just 10% included a conflict-of-interest statement. In short, transparency practices remain the exception rather than the norm, and they haven’t changed much between 2008/2009 and 2018/2019. Very briefly on what the authors did: they randomly sampled articles from Scopus at two time points—before and after the broader “replication crisis” was widely recognized—and descriptively coded them for openness indicators like open access, data and materials sharing, analysis code, preregistration, conflict-of-interest statements, and replication. Table 2 on page 8 summarizes the sample: out of 600 coded, 519 were included as linguistics studies, and 360 reported empirical work using either primary data they collected or secondary data such as corpora. Here’s what they found, in practical terms. Open access rose from 28% in 2008/2009 to 42% in 2018/2019, but still left over half the literature behind paywalls (section 3.2). Materials were available in about 16% of empirical studies overall, with a modest uptick for primary-data studies from 14.6% to 24.3% (section 3.3). Raw data sharing was rare for primary-data studies, dropping from 6.5% to 3.6%, but higher for secondary-data studies—34% to 42.6%—reflecting the reuse of public corpora (section 3.4.1). Processed data sharing nudged up from near-zero to single digits, and analysis scripts moved from 0% to about 1.5% by 2018/2019 (section 3.5). No articles reported preregistration (section 3.6). Only 1.1% of empirical studies were replications, all conceptual rather than direct (section 3.7). And conflict-of-interest statements appeared in just 10% of papers, usually to declare none (section 3.8). The bar chart in Figure 2 on page 10 makes this visually clear: across pre- and post-2018/2019, and for both primary and secondary data, the bars for materials, raw data, processed data, and analysis scripts stay low, with only small improvements. The paper also highlights a second, crucial imbalance: the field’s WEIRD skew—toward Western authors and English. Figure 1 on page 9 shows first-author affiliations concentrated in North America and Europe, and languages studied dominated by English; across all 600 articles, about 40.5% focused on English and roughly 60% on Indo-European languages (section 4.3). This limits how confidently we can generalize claims about “language” writ large. Why does this matter? Openness—open access, shared materials and data, analysis code, and preregistration—lets other researchers verify results, build on them efficiently, and run credible replications. Without it, errors are harder to catch, results are harder to reuse, and progress slows. And when most studies focus on a narrow slice of languages and researchers, we risk building theories that don’t travel well. The key takeaway: Linguistics has begun to move, but slowly. The authors argue for simple, concrete steps—share preprints and postprints; make materials, data, and code available when possible; preregister or consider Registered Reports; add conflict-of-interest statements; and value replications. Table 3 on page 21 lays out practical recommendations for researchers, journals, and institutions. The message is clear: more openness and broader linguistic and geographic diversity will make language science more trustworthy, useful, and truly global.

Today we’re looking at a health check on how open and reproducible linguistics research really is. The main finding is stark: across a random sample of 600 linguistics journal articles spanning two eras, only about a third were openly accessible, fewer than one in ten shared their materials, data, or analysis scripts, no articles were preregistered, only 1% were replications, and just 10% included a conflict-of-interest statement. In short, transparency practices remain the exception rather than the norm, and they haven’t changed much between 2008/2009 and 2018/2019.

Very briefly on what the authors did: they randomly sampled articles from Scopus at two time points—before and after the broader “replication crisis” was widely recognized—and descriptively coded them for openness indicators like open access, data and materials sharing, analysis code, preregistration, conflict-of-interest statements, and replication. Table 2 on page 8 summarizes the sample: out of 600 coded, 519 were included as linguistics studies, and 360 reported empirical work using either primary data they collected or secondary data such as corpora.

Here’s what they found, in practical terms. Open access rose from 28% in 2008/2009 to 42% in 2018/2019, but still left over half the literature behind paywalls (section 3.2). Materials were available in about 16% of empirical studies overall, with a modest uptick for primary-data studies from 14.6% to 24.3% (section 3.3). Raw data sharing was rare for primary-data studies, dropping from 6.5% to 3.6%, but higher for secondary-data studies—34% to 42.6%—reflecting the reuse of public corpora (section 3.4.1). Processed data sharing nudged up from near-zero to single digits, and analysis scripts moved from 0% to about 1.5% by 2018/2019 (section 3.5). No articles reported preregistration (section 3.6). Only 1.1% of empirical studies were replications, all conceptual rather than direct (section 3.7). And conflict-of-interest statements appeared in just 10% of papers, usually to declare none (section 3.8). The bar chart in Figure 2 on page 10 makes this visually clear: across pre- and post-2018/2019, and for both primary and secondary data, the bars for materials, raw data, processed data, and analysis scripts stay low, with only small improvements.

The paper also highlights a second, crucial imbalance: the field’s WEIRD skew—toward Western authors and English. Figure 1 on page 9 shows first-author affiliations concentrated in North America and Europe, and languages studied dominated by English; across all 600 articles, about 40.5% focused on English and roughly 60% on Indo-European languages (section 4.3). This limits how confidently we can generalize claims about “language” writ large.

Why does this matter? Openness—open access, shared materials and data, analysis code, and preregistration—lets other researchers verify results, build on them efficiently, and run credible replications. Without it, errors are harder to catch, results are harder to reuse, and progress slows. And when most studies focus on a narrow slice of languages and researchers, we risk building theories that don’t travel well.

The key takeaway: Linguistics has begun to move, but slowly. The authors argue for simple, concrete steps—share preprints and postprints; make materials, data, and code available when possible; preregister or consider Registered Reports; add conflict-of-interest statements; and value replications. Table 3 on page 21 lays out practical recommendations for researchers, journals, and institutions. The message is clear: more openness and broader linguistic and geographic diversity will make language science more trustworthy, useful, and truly global.