- Home
- e-Journals
- International Journal of Corpus Linguistics
- Previous Issues
- Volume 30, Issue 2, 2025
International Journal of Corpus Linguistics - Volume 30, Issue 2, 2025
Volume 30, Issue 2, 2025
-
Reproducibility, replicability, and robustness in corpus linguistics
Author(s): Martin Schweinberger and Michael Haughpp.: 119–129 (11)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThis introduction to the special issue Reproducibility, Replicability, and Robustness in Corpus Linguistics calls for more transparent and robust research practices in the field. It situates the discussion within the broader replication crisis in the life and social sciences and explores its relevance for corpus linguistics. The article identifies key areas for improvement — data management, workflows, and reporting — and showcases tools and principles such as FAIR/CARE, version control, reproducible notebooks, and open repositories. It highlights how corpus linguistics can build on open science infrastructures to enhance methodological rigor. Practical challenges, including data sensitivity and skill gaps, are addressed with actionable strategies. The issue brings together contributions that clarify core terminology, test the robustness of established methods, and suggest concrete ways forward. Together, these articles offer conceptual and practical guidance for making corpus linguistic research more open, verifiable, and aligned with broader scientific standards.
-
Reproducibility, replicability, robustness, and generalizability in corpus linguistics
Author(s): Joseph Flanaganpp.: 130–149 (20)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractEstablishing the credibility of scientific research involves several related but significantly different concerns. One potential problem in surveying different approaches to these concerns is that of terminology, as some of the basic terms used in the discussion — reproducibility, replicability, robustness, and generalizability — are often used in inconsistent or contradictory ways. This paper proposes to resolve such confusion by providing a terminological framework for discussing what kind of confirmation is necessary for a scientific study to be deemed credible. A study is said to be ‘reproducible’ if we can obtain identical results by performing an identical analysis on identical data, ‘replicable’ if we can obtain consistent results using the same analysis on different data, ‘robust’ if we can obtain consistent results from identical data using a different analysis, and ‘generalizable’ if we can obtain consistent results from different data using a different analysis.
-
Achieving stability in corpus-based analysis of word types
Author(s): Jesse Egbert, Douglas Biber, Bethany Gray and Tove Larssonpp.: 150–170 (21)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractRank-ordered lists of word types are ubiquitous in corpus linguistics and applied linguistics. Word lists are commonly developed as aids for language teaching and learning, vocabulary testing, and language description. Yet, these lists are often produced and used without evaluation of their stability — or replicability — across corpus samples. Our primary objective in this paper is to describe the cumulative state of knowledge regarding the stability of corpus-based word type lists, focusing on three goals that motivate the creation and use of rank-ordered lists: identifying key lexical items for learning or teaching, assessing vocabulary size or knowledge, and identifying all items in a language domain. We show that word type lists are far less stable than researchers and practitioners often assume, although there is substantial variability in stability depending on the goals and methods behind list creation.
-
Reuse of social media data in corpus linguistics
Author(s): Mikko Laitinen and Paula Rautionahopp.: 171–194 (24)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThe use of very large social media datasets in corpus linguistics has obvious benefits. Such data represent a novel source of evidence when compared with structured digital text corpora. However, there is a clear need to assess critically how the effective reuse of data can be handled, how findings can be reproduced, and how results can be generalized. A relevant question concerns the presentation of data to ensure reproducibility and replicability. This article surveys the state-of-the-art of descriptions of data collection and methodological transparency in 30 studies that used Twitter/X as their data. The empirical section investigates how easy it would be to reproduce a study based on these descriptions. While we concentrate on evidence from one social media application, the discussion continues to a presentation of concrete steps that might be used to improve data management related to the reuse, discovery, and evaluation of social media data in general.
-
Evaluating a transparent and interpretable approach to stance detection using linguistic markers in social media data
Author(s): Maud Reveilhac and Gerold Schneiderpp.: 195–233 (39)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractOur study focuses on replicability, which entails researchers’ ability to achieve similar results to a prior study using identical methods but a different yet comparable dataset. We address the challenge of stance detection (determining whether a document is “favorable,” “against,” or “neutral” toward a target), building on prior research underscoring the value of linguistic markers as complementary features for sentiment detection that enable more accurate stance classification. We utilize the Stance in Replies and Quotes (SRQ) dataset, which contains annotated discussion-based responses. Employing a rule-based methodology that emphasizes linguistic features, we examine whether the classification accuracy remains within a similar error margin as observed in a previous study of another dataset. Consistency is a necessary condition for robustness and generalizability, ultimately enhancing trust in the methodology. The replication of the model and its adaptability to the new data context demonstrate that it is competitive compared to existing machine learning studies.
-
Reproducibility and transparency in interpretive corpus pragmatics
Author(s): Martin Schweinberger and Michael Haughpp.: 234–260 (27)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractIn this paper we extend the discussion about reproducibility in corpus linguistics from quantitative to qualitative corpus-based approaches and argue that concerns about reproducibility can be addressed in interpretive research paradigms like corpus pragmatics. We first suggest that in interpretive research traditions, transparency is more important than reproducibility. We then argue that interpretive research can be made more transparent and accessible by using notebooks to share analytical procedures. We support these claims through a case study in which we analyse responses to information-seeking utterance-final or questions in spoken Australian English data. We use a qualitative, discourse analytic approach to systematically examine examples of these utterances from selected corpora. We show how corpus linguistic research can draw on existing infrastructures and tools for ensuring transparency, reproducibility, and replicability of interpretive analyses of the pragmatic functions of linguistic tokens in situated contexts.
Volumes & issues
-
Volume 31 (2026)
-
Volume 30 (2025)
-
Volume 29 (2024)
-
Volume 28 (2023)
-
Volume 27 (2022)
-
Volume 26 (2021)
-
Volume 25 (2020)
-
Volume 24 (2019)
-
Volume 23 (2018)
-
Volume 22 (2017)
-
Volume 21 (2016)
-
Volume 20 (2015)
-
Volume 19 (2014)
-
Volume 18 (2013)
-
Volume 17 (2012)
-
Volume 16 (2011)
-
Volume 15 (2010)
-
Volume 14 (2009)
-
Volume 13 (2008)
-
Volume 12 (2007)
-
Volume 11 (2006)
-
Volume 10 (2005)
-
Volume 9 (2004)
-
Volume 8 (2003)
-
Volume 7 (2002)
-
Volume 6 (2001)
-
Volume 5 (2000)
-
Volume 4 (1999)
-
Volume 3 (1998)
-
Volume 2 (1997)
-
Volume 1 (1996)
Most Read This Month
-
-
The Spoken BNC2014
Author(s): Robbie Love, Claire Dembry, Andrew Hardie, Vaclav Brezina and Tony McEnery
-
- More Less