- Home
- e-Journals
- International Journal of Corpus Linguistics
- Previous Issues
- Volume 31, Issue 3, 2026
International Journal of Corpus Linguistics - Volume 31, Issue 3, 2026
Volume 31, Issue 3, 2026
-
Deduplicating corpora : Introducing the Document Similarity tool
Author(s): Monika Bednarek and Kelvin K. H. Leepp.: 299–332 (34)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractCorpora containing newspaper articles are widely used in corpus-based discourse research. However, these corpora often contain duplicate texts due to issues such as syndication, the existence of agency copy, and largely similar print and online versions of the same article. While there are several tools that have automated deduplication functions, only a limited number of these tools allow corpus compilers to inspect duplicates before removal. In this paper, we will introduce the Document Similarity tool, which was developed to automatically remove identical texts and then provides near-identical candidates for users to examine side-by-side, giving corpus compilers more control of the deduplication process. We will describe several experiments using corpora that have been deduplicated using the Document Similarity tool in order to demonstrate its utility for corpus creation in corpus-based discourse analysis of newspapers and potentially other text types.
-
Not all linguistic variation is equally predictable : Evidence from two case studies with large language models
Author(s): Felix Morger and Aleksandrs Berdicevskispp.: 333–366 (34)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractWe compare to what extent the choice of a variant can be predicted from language-internal factors for two linguistic variables: the English dative alternation and the omission of the infinitival marker att in a Swedish future-tense construction. Previous research has shown that for the dative alternation, near-ceiling performance can be achieved by fitting a regression model with manually selected predictors. Similar attempts for att-omission have been unsuccessful. To test whether the two variables differ in predictability or whether optimal methods have not been found for att-omission, we apply a large language model (LLM) to the same task. For the dative alternation, LLM and regression perform equally well. For att-omission, LLM outperforms regression, but still performs worse than for the dative alternation, thus suggesting that att-omission is inherently less predictable. We argue that LLMs can be useful for estimating how much the choice of a variant depends on language-internal factors.
-
Sensitivity of dispersion measures to distributional patterns and corpus design
Author(s): Lukas Sönning and Jesse Egbertpp.: 367–395 (29)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractRecent work has shown that dispersion measures respond to multiple features in the data: Juilland’s D varies systematically with the number of corpus parts, and all commonly used indices are affected by the frequency of an item. This study uses a simulation approach to provide further insights into the sensitivity of dispersion measures to differences in corpus design (number of texts, average text length, distribution of text lengths) and distributional milieu (frequency and evenness of distribution). Our results suggest that, within the settings covered by our analysis, the factors frequency and evenness of distribution have roughly the same impact, though there is some variation among measures. The average text length emerges as another feature that leaves its mark on the observed scores. Finally, we note that D2 exhibits the same weakness as D — it varies with the number of corpus parts that enter the analysis.
-
Data interference : Emojis, homoglyphs, and issues of data fidelity in corpora and their results
Author(s): Matteo Di Cristofaropp.: 396–425 (30)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractTokenisation is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring the reliability of qualitative approaches. This paper examines how discrepancies in tokenisation affect the representation of language data and the validity of analytical findings. Investigating the challenges posed by emojis and homoglyphs, the study highlights the necessity of pre-processing these elements to maintain corpus fidelity to the source data. The research presents methods for ensuring that digital texts are accurately represented in corpora, thereby supporting reliable linguistic analysis and facilitating the repeatability of linguistic interpretations. The findings emphasise the necessity of a detailed understanding of both linguistic and technical aspects involved in digital textual data to enhance the accuracy of corpus analysis, and have significant implications for both quantitative and qualitative approaches in corpus-based research.
-
Review of Kaunisto & Schilk (2024): Challenges in corpus linguistics: Rethinking corpus compilation and analysis
Author(s): Holly Bakerpp.: 426–430 (5)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:This article reviews Challenges in corpus linguistics: Rethinking corpus compilation and analysis
-
Review of Partington & Diegoli (2026): Lexical priming: Evolution, evaluation and applications in English and Japanese
Author(s): József Andorpp.: 431–437 (7)show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:This article reviews Lexical priming: Evolution, evaluation and applications in English and Japanese
Volumes & issues
-
Volume 31 (2026)
-
Volume 30 (2025)
-
Volume 29 (2024)
-
Volume 28 (2023)
-
Volume 27 (2022)
-
Volume 26 (2021)
-
Volume 25 (2020)
-
Volume 24 (2019)
-
Volume 23 (2018)
-
Volume 22 (2017)
-
Volume 21 (2016)
-
Volume 20 (2015)
-
Volume 19 (2014)
-
Volume 18 (2013)
-
Volume 17 (2012)
-
Volume 16 (2011)
-
Volume 15 (2010)
-
Volume 14 (2009)
-
Volume 13 (2008)
-
Volume 12 (2007)
-
Volume 11 (2006)
-
Volume 10 (2005)
-
Volume 9 (2004)
-
Volume 8 (2003)
-
Volume 7 (2002)
-
Volume 6 (2001)
-
Volume 5 (2000)
-
Volume 4 (1999)
-
Volume 3 (1998)
-
Volume 2 (1997)
-
Volume 1 (1996)
Most Read This Month
-
-
The Spoken BNC2014
Author(s): Robbie Love, Claire Dembry, Andrew Hardie, Vaclav Brezina and Tony McEnery
-
- More Less