1887
image of The crystallization of language over time

Abstract

In this article, we investigate Van der Horst’s (2013) hypothesis that Dutch (like other European languages) underwent a diachronic process of ‘crystallization’, i.e. tighter lexical organization, at the expense of freely combinatorial syntax, in the last centuries. Analysing the collocational association strength in lemma and part-of-speech trigrams using the Δ measure and entropy (), we find quantitative support for the idea that Dutch has crystallized in the period under investigation (1850–1999). Further enquiry into the diachrony of the lexicon by means of Kullback-Leibler Divergence () suggests that the reason might be what Baayen et al. (2017) have called the ‘Ecclesiastes Principle’, namely a lexical expansion that puts a cap on the combinatorial syntax. This lexical expansion is probably a response to the increasing specialization and cultural turnover in late modern times. The slow change in the syntax of Dutch shows how languages adapt to their cultural niche.

Available under the CC BY 4.0 license.
Loading

Article metrics loading...

/content/journals/10.1075/ijcl.24105.det
2026-08-18
2026-09-09
Loading full text...

Full text loading...

/deliver/fulltext/10.1075/ijcl.24105.det/ijcl.24105.det.html?itemId=/content/journals/10.1075/ijcl.24105.det&mimeType=html&fmt=ahah

References

  1. Arnon, I., & Snider, N.
    (2010) More than words: Frequency effects for multi-word phrases. Journal of Memory and Language, (), –. 10.1016/j.jml.2009.09.005
    https://doi.org/10.1016/j.jml.2009.09.005 [Google Scholar]
  2. Baayen, R. H., & Linke, M.
    (2020) Generalized additive mixed models. InM. Paquot & S. Th. Gries (Eds.), A Practical handbook of corpus linguistics (pp.–). Springer. 10.1007/978‑3‑030‑46216‑1_23
    https://doi.org/10.1007/978-3-030-46216-1_23 [Google Scholar]
  3. Baayen, R. H., Tomaschek, F., Gahl, S., & Ramscar, M.
    (2017) The Ecclesiastes Principle in language change. InM. Hundt, S. Mollin, & S. Pfenninger (Eds.) The changing English language: Psycholinguistic perspectives (pp.–). Cambridge University Press. 10.1017/9781316091746.002
    https://doi.org/10.1017/9781316091746.002 [Google Scholar]
  4. Barlow, M.
    (2013) Individual differences and usage-based grammar. International Journal of Corpus Linguistics, (), –. 10.1075/ijcl.18.4.01bar
    https://doi.org/10.1075/ijcl.18.4.01bar [Google Scholar]
  5. Bentz, C., Alikaniotis, D., Cysouw, M., & Ferrer-i-Cancho, R.
    (2017) The Entropy of words: Learnability and expressivity across more than 1000 languages. Entropy, (), . 10.3390/e19060275
    https://doi.org/10.3390/e19060275 [Google Scholar]
  6. Bentz, C., & Berdicevskis, A.
    (2016) Learning pressures reduce morphological complexity: Linking corpus, computational and experimental evidence. InD. Brumato, F. Dell’Orletta, G. Venturi, T. François, & P. Blache (Eds.), Proceedings of the workshop on computational linguistics for linguistic complexity (pp.–). The COLING 2016 Organizing Committee. https://aclanthology.org/W16-4125.pdf
    [Google Scholar]
  7. Bentz, C., & Winter, B.
    (2013) Languages with more second language learners tend to lose nominal case. Language Dynamics and Change, (), –. 10.1163/22105832‑13030105
    https://doi.org/10.1163/22105832-13030105 [Google Scholar]
  8. Bizzoni, Y., Degaetano-Ortlieb, S., Fankhauser, P., & Teich, E.
    (2020) Linguistic variation and change in 250 years of English scientific writing: A data-driven approach. Frontiers in Artificial Intelligence, , . 10.3389/frai.2020.00073
    https://doi.org/10.3389/frai.2020.00073 [Google Scholar]
  9. Chen, S., Gil, D., Gaponov, S., Reifegerste, J., Yuditha, T., Tatarinova, T., Progovac, L., & Benítez-Burraco, A.
    (2024) Linguistic correlates of societal variation: A quantitative analysis. PLoS ONE, (), . 10.1371/journal.pone.0300838
    https://doi.org/10.1371/journal.pone.0300838 [Google Scholar]
  10. Conklin, K., & Schmitt, N.
    (2008) Formulaic sequences: Are they processed more quickly than nonformulaic language by native and nonnative speakers?Applied Linguistics, (), –. 10.1093/applin/amm022
    https://doi.org/10.1093/applin/amm022 [Google Scholar]
  11. Coupé, C.
    (2018) Modeling linguistic variables with regression models: Addressing non-Gaussian distributions, non-independent observations, and non-linear predictors with random effects and generalized additive Models for location, scale, and shape. Frontiers in Psychology, , . 10.3389/fpsyg.2018.00513
    https://doi.org/10.3389/fpsyg.2018.00513 [Google Scholar]
  12. Dąbrowska, E.
    (2014) Recycling utterances: A speaker’s guide to sentence processing. Cognitive Linguistics, (), –. 10.1515/cog‑2014‑0057
    https://doi.org/10.1515/cog-2014-0057 [Google Scholar]
  13. Daelemans, W., & Van den Bosch, A.
    (2005) Memory-based language processing. Cambridge University Press. 10.1017/CBO9780511486579
    https://doi.org/10.1017/CBO9780511486579 [Google Scholar]
  14. Degaetano-Ortlieb, S., & Teich, E.
    (2022) Toward an optimal code for communication: The case of scientific English. Corpus Linguistics and Linguistic Theory, (), –. 10.1515/cllt‑2018‑0088
    https://doi.org/10.1515/cllt-2018-0088 [Google Scholar]
  15. De Troij, R.
    (2023) Natiolectal variation in Dutch grammar: A data-driven approach. [PhD Dissertation]. KU Leuven & Radboud University Nijmegen. https://research.kuleuven.be/portal/en/project/3H170584
  16. Dowle, M., & Srinivasan, A.
    (2024) data.table: Extension of ‘data.frame’ (Version 1.16.4). R package. https://prism.dev.a2-ai.cloud/docs/data.table/1.16.4/?isIframe=true&nocache=1775435972
    [Google Scholar]
  17. Dunn, J.
    (2018) Multi-unit association measures: Moving beyond pairs of words. International Journal of Corpus Linguistics, (), –. 10.1075/ijcl.16098.dun
    https://doi.org/10.1075/ijcl.16098.dun [Google Scholar]
  18. Ellis, N. C.
    (2006) Language acquisition as rational contingency learning. Applied Linguistics, (), –. 10.1093/applin/ami038
    https://doi.org/10.1093/applin/ami038 [Google Scholar]
  19. Evert, S.
    (2009) Corpora and collocations. InA. Lüdeling & M. Kytö (Eds.), Corpus linguistics: An international handbook Vol. 2 (pp.–). De Gruyter Mouton. 10.1515/9783110213881.2.1212
    https://doi.org/10.1515/9783110213881.2.1212 [Google Scholar]
  20. Fankhauser, P., Knappen, J., & Teich, E.
    (2014) Exploring and visualizing variation in language resources. InN. Calzolari, K. Choukri, T. Declerck, H. Loftsson, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, & S. Piperidis (Eds.) Proceedings of the ninth international conference on Language Resources and Evaluation (LREC’14) (pp.–). European Language Resources Association. www.lrec-conf.org/proceedings/lrec2014/pdf/185_Paper.pdf. 10.63317/3txgs8z92rrr
    https://doi.org/10.63317/3txgs8z92rrr [Google Scholar]
  21. Fox, J., & Weisberg, S.
    (2019) An R companion to applied regression (3rd ed.). Thousand Oaks, CA: SAGE Publications, Inc.
    [Google Scholar]
  22. Goldberg, A. E.
    (1995) Constructions: A construction grammar approach to argument structure. University of Chicago Press.
    [Google Scholar]
  23. (2006) Constructions at work: The nature of generalization in language. Oxford University Press. 10.1093/acprof:oso/9780199268511.001.0001
    https://doi.org/10.1093/acprof:oso/9780199268511.001.0001 [Google Scholar]
  24. Gries, S. Th.
    (2015) 15-something years of work on collocations: What is or should be next…International Journal of Corpus Linguistics, (), –. 10.1075/ijcl.18.1.09gri
    https://doi.org/10.1075/ijcl.18.1.09gri [Google Scholar]
  25. Grondelaers, S., Speelman, D., & Geeraerts, D.
    (2008) National variation in the use of er ‘there’: Regional and diachronic constraints on cognitive explanations. InG. Kristiansen & R. Dirven (Eds.), Cognitive sociolinguistics: Language variation, cultural models, social systems (pp.–). De Gruyter Mouton. 10.1515/9783110199154.2.153
    https://doi.org/10.1515/9783110199154.2.153 [Google Scholar]
  26. Ham, L., Lentz, L., Pander Maat, H., & Stolk, F.
    (2018) Zijn romans en kranten sinds 1950 eenvoudiger geworden?Tijdschrift voor Nederlandse Taal- en Letterkunde, (), –. https://www.tntl.nl/index.php/tntl/article/view/531
    [Google Scholar]
  27. Helasvuo, M.-L.
    (2014) Agreement or crystallization: Patterns of 1st and 2nd person subjects and verbs of cognition in Finnish conversational interaction. Journal of Pragmatics, , –. 10.1016/j.pragma.2013.11.011
    https://doi.org/10.1016/j.pragma.2013.11.011 [Google Scholar]
  28. Hoffmann, T., & Trousdale, G.
    (Eds.) (2013) The Oxford handbook of construction grammar. Oxford University Press. 10.1093/oxfordhb/9780195396683.001.0001
    https://doi.org/10.1093/oxfordhb/9780195396683.001.0001 [Google Scholar]
  29. Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A.
    (2020) spaCy: Industrial-strength natural language processing in Python. 10.5281/zenodo.1212303
    https://doi.org/10.5281/zenodo.1212303 [Google Scholar]
  30. Jurafsky, D., & Martin, J. H.
    (2026) Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition with language models (3rd edn). https://web.stanford.edu/~jurafsky/slp3
    [Google Scholar]
  31. Koplenig, A., & Wolfer, S.
    (2023) Languages with more speakers tend to be harder to (machine-)learn. Scientific Reports, (), . 10.1038/s41598‑023‑45373‑z
    https://doi.org/10.1038/s41598-023-45373-z [Google Scholar]
  32. Kusters, W.
    (2003) Linguistic complexity: The influence of social change on verbal inflection. LOT. https://www.lotpublications.nl/Documents/077_fulltext.pdf
    [Google Scholar]
  33. Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B.
    (2017) lmerTest package: Tests in linear mixed effects models. Journal of Statistical Software, (), –. 10.18637/jss.v082.i13
    https://doi.org/10.18637/jss.v082.i13 [Google Scholar]
  34. Lev-Ari, S.
    (2024) The influence of community structure on how communities categorize the world. Journal of Experimental Psychology: Learning, Memory, and Cognition, (), –. 10.1037/xlm0001334
    https://doi.org/10.1037/xlm0001334 [Google Scholar]
  35. Lupyan, G., & Dale, R.
    (2010) Language structure is partly determined by social structure. PLoS ONE, (), . 10.1371/journal.pone.0008559
    https://doi.org/10.1371/journal.pone.0008559 [Google Scholar]
  36. Meillet, A.
    (1922) Caractères généraux des langues germaniques. Librairie Hachette.
    [Google Scholar]
  37. Mohseni, M., Redies, C., & Gast, V.
    (2023) Comparative analysis of preference in contemporary and earlier texts using entropy measures. Entropy, (), . 10.3390/e25030486
    https://doi.org/10.3390/e25030486 [Google Scholar]
  38. Müller, K.
    (2020) here: A simpler way to find your files (Version 1.0.1). R package. https://here.r-lib.org/
    [Google Scholar]
  39. Nijs, J., Van de Velde, F., & Cuyckens, H.
    (2025) Is word order responsive to morphology? Using Kolmogorov complexity and Granger causality to disentangle cause and effect in morphosyntactic change in five Western European languages. Entropy, (), . 10.3390/e27010053
    https://doi.org/10.3390/e27010053 [Google Scholar]
  40. Pedersen, T. L.
    (2024) Patchwork: The composer of plots (Version 1.3.0). R package. https://patchwork.data-imaginist.com/
    [Google Scholar]
  41. Petré, P., & Van de Velde, F.
    (2018) The real-time dynamics of the individual and the community in grammaticalization. Language, (), –. 10.1353/lan.2018.0056
    https://doi.org/10.1353/lan.2018.0056 [Google Scholar]
  42. Piersoul, J., De Troij, R., & Van de Velde, F.
    (2021) 150 years of written Dutch: The construction of the Dutch Corpus of Contemporary and Late Modern Periodicals. Nederlandse Taalkunde, (), –. 10.5117/NEDTAA2021.3.002.PIER
    https://doi.org/10.5117/NEDTAA2021.3.002.PIER [Google Scholar]
  43. Pijpops, D.
    (2021) Does standardization affect the type of motivating factors that determine language variation? The case of the Dutch transitive-reflexive alternation. InG. Kristiansen, K. Franco, S. De Pascale, L. Rosseel, & W. Zhang (Eds.), Cognitive sociolinguistics revisited (pp.–). De Gruyter Mouton. 10.1515/9783110733945‑030
    https://doi.org/10.1515/9783110733945-030 [Google Scholar]
  44. Piotrowski, M.
    (2012) Natural language processing for historical texts. Springer. 10.1007/978‑3‑031‑02146‑6
    https://doi.org/10.1007/978-3-031-02146-6 [Google Scholar]
  45. R Core Team
    R Core Team (2024) R: A language and environment for statistical computing [Computer software]. R Foundation for Statistical Computing. https://www.R-project.org/
    [Google Scholar]
  46. Raviv, L., Meyer, A., & Lev-Ari, S.
    (2019) Larger communities create more systematic languages. Proceedings of the Royal Society B — Biological Sciences, (), . 10.1098/rspb.2019.1262
    https://doi.org/10.1098/rspb.2019.1262 [Google Scholar]
  47. Ryckaert, R.
    (2017) Ruggespraak in het gekkenhuis: Een beknopte geschiedenis van de spelling van het Nederlands. InG. De Sutter (ed.), De vele gezichten van het Nederlands in Vlaanderen: Een inleiding tot de variatietaalkunde (pp.–). Acco.
    [Google Scholar]
  48. Sapir, E.
    (1921) Language: An introduction to the study of speech. Harcourt.
    [Google Scholar]
  49. Shcherbakova, O., Michaelis, S. M., Haynie, H. J., Passmore, S., Gast, V., Gray, R. D., Greenhill, S. J., Blasi, D. E., & Skirgård, H.
    (2023) Societies of strangers do not speak less complex languages. Science Advances, (), . 10.1126/sciadv.adf7704
    https://doi.org/10.1126/sciadv.adf7704 [Google Scholar]
  50. Speelman, D., Grondelaers, S., Szmrecsanyi, B., & Heylen, K.
    (2020) Schaalvergroting in het syntactische alternantieonderzoek: Een nieuwe analyse van het presentatieve er met automatisch gegenereerde predictoren. Nederlandse Taalkunde, (), –. 10.5117/NEDTAA2020.1.005.SPEE
    https://doi.org/10.5117/NEDTAA2020.1.005.SPEE [Google Scholar]
  51. Tremblay, A., Derwing, B., Libben, G., & Westbury, C.
    (2011) Processing advantages of lexical bundles: Evidence from self-paced reading and sentence recall tasks. Language Learning, (), –. 10.1111/j.1467‑9922.2010.00622.x
    https://doi.org/10.1111/j.1467-9922.2010.00622.x [Google Scholar]
  52. Trudgill, P.
    (2011) Sociolinguistic typology: Social determinants of linguistic complexity. Oxford University Press.
    [Google Scholar]
  53. Van Bree, C.
    (1987) Historische grammatica van het Nederlands. Foris.
    [Google Scholar]
  54. Van Cranenburgh, A., & Van Noord, G.
    (2022) OpenBoek: A corpus of literary coreference and entities with an exploration of historical spelling normalization. Computational Linguistics in the Netherlands Journal, , –. https://clinjournal.org/clinj/article/view/157
    [Google Scholar]
  55. Van der Horst, J.
    (1997) Over en naar aanleiding van Zuid-Nederlandse doorbrekingen. InA. Van Santen & M. J. Van der Wal (Eds.), Taal in tijd en ruimte: Voor Cor van Bree bij zijn afscheid als hoogleraar Historische Taalkunde en Taalvariatie aan de Vakgroep Nederlands van de Rijksuniversiteit Leiden (pp.–). Stichting Neerlandistiek.
    [Google Scholar]
  56. (2008) Het einde van de standaardtaal: Een wisseling van Europese taalcultuur. Meulenhoff.
    [Google Scholar]
  57. (2013) Taal op drift: Lange-termijnontwikkelingen in taal en samenleving. Meulenhoff.
    [Google Scholar]
  58. Van de Velde, F.
    (2009a) The emergence of modification patterns in the Dutch noun phrase. Linguistics, (), –. 10.1515/LING.2009.036
    https://doi.org/10.1515/LING.2009.036 [Google Scholar]
  59. (2009b) De nominale constituent: Structuur en geschiedenis. Leuven University Press.
    [Google Scholar]
  60. (2010) The emergence of the determiner in the Dutch NP. Linguistics, (), –. 10.1515/ling.2010.009
    https://doi.org/10.1515/ling.2010.009 [Google Scholar]
  61. Van de Velde, F., Piersoul, J., De Smet, I., & Ruppert, E.
    (2021) Changing preferences in cultural references. InG. Kristiansen, K. Franco, S. De Pascale, L. Rosseel, & W. Zhang (Eds.), Cognitive sociolinguistics revisited (pp.–). De Gruyter Mouton. 10.1515/9783110733945‑047
    https://doi.org/10.1515/9783110733945-047 [Google Scholar]
  62. Van Eynde, F.
    (2004) Part of speech tagging and lemmatizing of the Corpus Gesproken Nederlands (Spoken Dutch Corpus). Centrum voor Computerlinguïstiek K. U. Leuven. https://ivdnt.org/images/stories/producten/documentatie/cgn_website/doc_English/topics/annot/pos_tagging/tg_prot_en.pdf
    [Google Scholar]
  63. Van Rij, J., Wieling, M., Baayen, R. H., & Van Rijn, D.
    (2022) itsadug: Interpreting time series and autocorrelated data using GAMMs (Version 2.4.1). R package. https://rdrr.io/cran/itsadug/man/itsadug.html
    [Google Scholar]
  64. Wickham, H., François, R., Henry, L., Müller, K., & Vaughan, D.
    (2023) dplyr: A grammar of data manipulation (Version 1.1.4). R package. https://cran.r-project.org/web/packages/dplyr/index.html
    [Google Scholar]
  65. Wickham, H., Vaughan, D., & Girlich, M.
    (2024) tidyr: tidy messy data (Version 1.3.1). R package. https://tidyr.tidyverse.org/
    [Google Scholar]
  66. Wickham, H.
    (2016) ggplot2: Elegant graphics for data analysis (2nd ed.). Springer. 10.1007/978‑3‑319‑24277‑4
    https://doi.org/10.1007/978-3-319-24277-4 [Google Scholar]
  67. Wood, S. N.
    (2003) Thin-plate regression splines. Journal of the Royal Statistical Society (B), (). –. 10.1111/1467‑9868.00374
    https://doi.org/10.1111/1467-9868.00374 [Google Scholar]
  68. (2011) Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. Journal of the Royal Statistical Society (B), (), –. 10.1111/j.1467‑9868.2010.00749.x
    https://doi.org/10.1111/j.1467-9868.2010.00749.x [Google Scholar]
  69. (2017) Generalized additive models: An introduction with R (2nd ed.). Chapman and Hall/CRC. 10.1201/9781315370279
    https://doi.org/10.1201/9781315370279 [Google Scholar]
/content/journals/10.1075/ijcl.24105.det
Loading
/content/journals/10.1075/ijcl.24105.det
Loading

Data & Media loading...

  • Article Type: Research Article
Keywords: ΔP ; entropy ; linguistic niche hypothesis ; crystallization ; Dutch
This is a required field
Please enter a valid email address
Approval was successful
Invalid data
An Error Occurred
Approval was partially successful, following selected items could not be processed due to error