- Home
- e-Journals
- International Journal of Corpus Linguistics
- Fast Track Listing
International Journal of Corpus Linguistics - Online First
Online First articles are the published Version of Record, made available as soon as they are finalized and formatted. They are in general accessible to current subscribers, until they have been included in an issue, which is accessible to subscribers to the relevant volume
-
-
The passive voice of subjective experience and objective response : A corpus-assisted analysis of the semantic–pragmatic interface of the Chinese 遭 zāo, 受 shòu, and 得到 dédào constructions
Author(s): Laura LocatelliAvailable online: 10 September 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThe passive voice in Chinese is the focus of extensive research seeking to understand its diverse discourse functions. Adopting a corpus-assisted approach, this paper brings attention to an underexplored category of passives — Chinese lexical passive constructions (LPCs) — and analyzes the semantic and pragmatic aspects of the zāo, shòu, and dédào structures. Quantitative and qualitative analysis helps to identify the semantic frames triggered by each LPC, highlighting potential differences in how events are conceptualized. Based on evidence from three corpora representing distinct discourse domains, the paper contends that LPCs differ in the range of events they represent as well as the affective valence they convey. As such, they constitute three distinct ways of conceptualizing the effects of an action, allowing the speaker to express their subjective-objective response to the event. This supports the unique form-function relationship of each LPC and informs the pragmatic implications involved in their selection.
-
-
-
dassBNC2014 : Retrieving and mapping directive speech acts in an open-domain corpus of everyday conversations
Author(s): Guying Zhou and Jiajin XuAvailable online: 07 September 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThis study introduces dassBNC2014, a manually annotated subset of the Spoken British National Corpus 2014 designed to capture directive speech acts in everyday conversation. In natural language processing (NLP), speech act annotations in open-domain dialog remain scarce, while manually annotated data are fragmented and inconsistently accessible. With dassBNC2014, we aim to address the persistent gaps in existing resources. Our annotation scheme combines the ISO 24617-2 framework with insights from speech act theory to achieve an intermediate level of granularity, which is both theoretically grounded and operationally feasible. Analysis of the annotated directives shows that they constitute a heterogeneous set of practices varying systematically in interactional locus, temporal orientation, deontic strength and action type. These findings refine the theoretical description of directive force, establishing dassBNC2014 as a resource for advancing corpus-based speech act research and supporting the development of NLP models capable of recognizing directive speech acts in authentic conversational data.
-
-
-
Accounting for register-internal variation : Toward a framework for analyzing communicative purpose and register comparisons for functional correspondence
Author(s): Marianna Gracheva and Jesse EgbertAvailable online: 04 September 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractRegisters are text varieties associated with situations of use. It has now been documented, however, that texts within registers are not situationally homogeneous — just as they are not linguistically uniform. This led Biber and Egbert (2023) to propose that linguistic variation within registers functionally corresponds to situational variation among their texts. The situational factor consistently shown to vary within registers is communicative purpose. However, previous studies have examined variability in purpose within a single register and always relied on coding schemes developed for that register. As a result, there is no single taxonomy of communicative purposes to be applied across registers. This paper proposes such a framework, offering guidance for its future adaptations to other corpora, and demonstrates how its application enables (a) analyses of functional correspondence between communicative and linguistic variation among texts of several registers; (b) register comparisons for the extent of internal variation.
-
-
-
The crystallization of language over time
Author(s): Robbert De Troij and Freek Van de VeldeAvailable online: 18 August 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractIn this article, we investigate Van der Horst’s (2013) hypothesis that Dutch (like other European languages) underwent a diachronic process of ‘crystallization’, i.e. tighter lexical organization, at the expense of freely combinatorial syntax, in the last centuries. Analysing the collocational association strength in lemma and part-of-speech trigrams using the ΔP measure and entropy (H), we find quantitative support for the idea that Dutch has crystallized in the period under investigation (1850–1999). Further enquiry into the diachrony of the lexicon by means of Kullback-Leibler Divergence (KLD) suggests that the reason might be what Baayen et al. (2017) have called the ‘Ecclesiastes Principle’, namely a lexical expansion that puts a cap on the combinatorial syntax. This lexical expansion is probably a response to the increasing specialization and cultural turnover in late modern times. The slow change in the syntax of Dutch shows how languages adapt to their cultural niche.
-
-
-
“There’s a lot more Chinese people” : Tracking ethnic variation in plural there existentials over time in Australian English
Author(s): Qiao GanAvailable online: 13 August 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThe variable realisation of there+be+plural arguments (e.g. there’s a lot more Chinese people) is well-documented across English varieties, but little research has examined its use among ethnic minorities or how intersecting social factors shape its variation. This study addresses these gaps, analysing 1,549 tokens from 206 Australians of Anglo, Italian, Greek, and Chinese backgrounds, recorded in the 1970s and 2010s. The study finds an apparent-time increase in there’s during the 1970s, led by teenagers across ethnic groups, followed by broad adoption by the 2010s. Usage patterns reflect intersections of age, ethnicity, and class, with working-class Adult Anglos consistently favouring there’s. Additionally, the grammatical conditioning of (there’s) has shifted: the definiteness effect, where definite and numerical determiners (e.g. these, two) favour there’s, emerged among 1970s teenagers and has since been maintained. Changes in frequency are accompanied by shifts in linguistic and social conditioning, illustrating how variation evolves within complex sociolinguistic ecologies.
-
-
-
Introducing HUM19 : A corpus of 19th century British and Irish fiction
Author(s): Fransina Stradling, Brian Walker, Hazel Price, Dan McIntyre and Michael BurkeAvailable online: 04 August 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThis report introduces the HUM19 corpus of 19th century British and Irish novels and describes its construction. We characterise the HUM19 corpus as a specialised reference corpus that can be used both as a comparator against which other corpora or texts may be compared and as a study corpus, searchable for the linguistic features and patterns associated with 19th century British and Irish fiction. We explain how these envisioned uses informed the construction process and show that the construction of the HUM19 corpus required principled compromise with regard to sampling but that this was achieved without necessarily forfeiting corpus quality.
-
-
-
From unannotated to annotated : The impact of text internal variation on register predictions of long historical documents
Author(s): Liina Repo, Brett Hashimoto and Veronika LaippalaAvailable online: 10 June 2026show More to view fulltext, buy and share links for: show Less to hide fulltext, buy and share links for:AbstractThe utility of historical language databases is hindered by their complexity, text variety, and lack of register information. We investigate the feasibility of deriving linguistically motivated and reliable register predictions for unannotated, long historical texts. We fine-tune BERT-based deep learning models using register-annotated data from the Corpus of Founding Era American English and predict registers for different text parts of unannotated Eighteenth Century English Online documents. We determine the model’s effectiveness in capturing pervasive linguistic features across different text sections (e.g. beginnings vs. endings), analyzing how document internal variation affects model performance and identifying which sections are most effectively predicted. Additionally, we employ the Stable Attribution Class Explanation method to extract and compare keywords from various text parts to determine the quality of the predictions. Our findings indicate that text beginnings consistently provide more reliable classifications.
-
Most Read This Month Most Read RSS feed
-
-
The Spoken BNC2014
Author(s): Robbie Love, Claire Dembry, Andrew Hardie, Vaclav Brezina and Tony McEnery
-
- More Less