1887
image of Deduplicating corpora

Abstract

Corpora containing newspaper articles are widely used in corpus-based discourse research. However, these corpora often contain duplicate texts due to issues such as syndication, the existence of agency copy, and largely similar print and online versions of the same article. While there are several tools that have automated deduplication functions, only a limited number of these tools allow corpus compilers to inspect duplicates before removal. In this paper, we will introduce the Document Similarity tool, which was developed to automatically remove identical texts and then provides near-identical candidates for users to examine side-by-side, giving corpus compilers more control of the deduplication process. We will describe several experiments using corpora that have been deduplicated using the Document Similarity tool in order to demonstrate its utility for corpus creation in corpus-based discourse analysis of newspapers and potentially other text types.

Available under the CC BY 4.0 license.
Loading

Article metrics loading...

/content/journals/10.1075/ijcl.24197.bed
2026-06-16
2026-07-17
Loading full text...

Full text loading...

/deliver/fulltext/10.1075/ijcl.24197.bed/ijcl.24197.bed.html?itemId=/content/journals/10.1075/ijcl.24197.bed&mimeType=html&fmt=ahah

References

  1. Anthony, L.
    (2024) AntConc (Version 4.3.1) [Computer software]. Waseda University. https://www.laurenceanthony.net/software/antconc/
    [Google Scholar]
  2. Baker, P.
    (2023) Using corpora in discourse analysis (2nd ed.). Bloomsbury Academic. 10.5040/9781350083783
    https://doi.org/10.5040/9781350083783 [Google Scholar]
  3. Balfour, J.
    (2023) Representing schizophrenia in the media: A corpus-based approach to UK press coverage. Routledge. 10.4324/9781003096054
    https://doi.org/10.4324/9781003096054 [Google Scholar]
  4. Bednarek, M.
    (2016) Voices and values in the news: News media talk, news values and attribution. Discourse, Context & Media, , –. 10.1016/j.dcm.2015.11.004
    https://doi.org/10.1016/j.dcm.2015.11.004 [Google Scholar]
  5. Bednarek, M., Bray, C., Coltman-Patel, T., & Bonfiglioli, C.
    (2026) Examining the uptake of media guidelines: A corpus analysis of obesity representation in Australian and UK news. InG. Brookes, N. Curry, & R. Love (Eds.), Applications of corpus linguistics: Established and emergent contexts (pp.–). Cambridge University Press. 10.1017/9781009382007.010
    https://doi.org/10.1017/9781009382007.010 [Google Scholar]
  6. Bednarek, M., Bray, C., Vanichkina, D. P., Brookes, G., Bonfiglioli, C., Coltman-Patel, T., Lee, K., & Baker, P.
    (2024) Weight stigma: Towards a language-informed analytical framework. Applied Linguistics, (), –. 10.1093/applin/amad033
    https://doi.org/10.1093/applin/amad033 [Google Scholar]
  7. Bednarek, M., & Caple, H.
    (2017) The Discourse of news values: How news organizations create newsworthiness. Oxford University Press. 10.1093/acprof:oso/9780190653934.001.0001
    https://doi.org/10.1093/acprof:oso/9780190653934.001.0001 [Google Scholar]
  8. Bednarek, M., & Carr, G.
    (2019) Guide to the Diabetes News Corpus (DNC). Available at: https://osf.io/jrhx2/
    [Google Scholar]
  9. (2020) Diabetes coverage in Australian newspapers (2013–2017): A computer-based linguistic analysis. Health Promotion Journal of Australia, (), –. 10.1002/hpja.295
    https://doi.org/10.1002/hpja.295 [Google Scholar]
  10. (2021a) Computer-assisted digital text analysis for journalism and communications research: Introducing corpus linguistic techniques that do not require programming. Media International Australia, (), –. 10.1177/1329878X20947124
    https://doi.org/10.1177/1329878X20947124 [Google Scholar]
  11. (2021b) Australian diabetes news media coverage. Australian Diabetes Educator, (). https://ade.adea.com.au/australian-diabetes-news-media-coverage/
    [Google Scholar]
  12. Bednarek, M., Schweinberger, M., & Lee, K. K. H.
    (2024) Corpus-based discourse analysis: From meta-reflection to accountability. Corpus Linguistics and Linguistic Theory, (), –. 10.1515/cllt‑2023‑0104
    https://doi.org/10.1515/cllt-2023-0104 [Google Scholar]
  13. Bell, A.
    (1991) The language of news media. Blackwell.
    [Google Scholar]
  14. Brezina, V., Hawtin, A., & McEnery, T.
    (2021) The Written British National Corpus 2014 — Design and comparability. Text & Talk, (), –. 10.1515/text‑2020‑0052
    https://doi.org/10.1515/text-2020-0052 [Google Scholar]
  15. Broder, A. Z.
    (1998) On the resemblance and containment of documents. InB. Carpentieri, A. De Santis, U. Vaccaro, & J. A. Storer (Eds.), Proceedings: Compression and complexity of SEQUENCES 1997 (pp.–). Institute of Electrical and Electronics Engineers. 10.1109/SEQUEN.1997.666900
    https://doi.org/10.1109/SEQUEN.1997.666900 [Google Scholar]
  16. Brookes, G., & Baker, P.
    (2021) Obesity in the news: Language and representation in the press. Cambridge University Press. 10.1017/9781108864732
    https://doi.org/10.1017/9781108864732 [Google Scholar]
  17. Coltman-Patel, T.
    (2020) Weight stigma in Britain: The linguistic representation of obesity in newspapers [Doctoral dissertation, Nottingham Trent University]. Nottingham Trent University’s Institutional Repository (IRep). irep.ntu.ac.uk/id/eprint/42065/
  18. (2023) (Mis)Representing weight and obesity in the British press: Fear, divisiveness, shame and stigma. Palgrave Macmillan. 10.1007/978‑3‑031‑44854‑6
    https://doi.org/10.1007/978-3-031-44854-6 [Google Scholar]
  19. Fairclough, N.
    (1995) Media discourse. Hodder Arnold.
    [Google Scholar]
  20. Foley, K., Ward, P., & McNaughton, D.
    (2019) Innovating qualitative framing analysis for purposes of media analysis within public health inquiry. Qualitative Health Research, (), –. 10.1177/1049732319826559
    https://doi.org/10.1177/1049732319826559 [Google Scholar]
  21. Fuoli, M., & Bednarek, M.
    (2022) Emotional labor in webcare and beyond: A linguistic framework and case study. Journal of Pragmatics, , –. 10.1016/j.pragma.2022.01.016
    https://doi.org/10.1016/j.pragma.2022.01.016 [Google Scholar]
  22. Grant, S., Soltani Panah, A., & McCosker, A.
    (2022) Weight-biased language across 30 years of Australian news reporting on obesity: Associations with public health policy. Obesities, (), –. 10.3390/obesities2010010
    https://doi.org/10.3390/obesities2010010 [Google Scholar]
  23. Hochschild, A. R.
    (1983) The Managed heart: Commercialization of human feeling. University of California Press.
    [Google Scholar]
  24. Jufri, S., & Sun, C.
    (2024) Document Similarity (v1.2.0) [Computer software]. Language Data Commons of Australia (Australian Text Analytics Platform). 10.5281/zenodo.20542430
    https://doi.org/10.5281/zenodo.20542430 [Google Scholar]
  25. Kilgarriff, A., Rychlý, P., Smrž, P., & Tugwell, D.
    (2004) The Sketch Engine. InG. Williams & S. Vessier (Eds.), Proceedings of the 11th EURALEX international congress (pp.–). European Association for Lexicography. https://euralex.org/publications/the-sketch-engine/
    [Google Scholar]
  26. Leitner, G.
    (1986) Reporting the ‘events of the day’: Uses and functions of reported speech. Studia Anglica Posnaniensia, , –.
    [Google Scholar]
  27. LexisNexis
    LexisNexis. (n.d.-a). Deduplication in newsdesk searches and newsletters: Document ID HT7156. LexisNexis Support Center. https://supportcenter.lexisnexis.com/app/answers/answer_view/a_id/1100627/loc/en_US
    [Google Scholar]
  28. LexisNexis
    LexisNexis. (n.d.-b). Group duplicates: Document ID HT5459. LexisNexis Support Center. https://supportcenter.lexisnexis.com/app/answers/answer_view/a_id/1089709/loc/en_US
    [Google Scholar]
  29. Mockler, N.
    (2022) Constructing teacher identities: How the print media define and represent teachers and their work. Bloomsbury Academic. 10.5040/9781350132917
    https://doi.org/10.5040/9781350132917 [Google Scholar]
  30. Price, H.
    (2022) The Language of mental illness: Corpus linguistics and the construction of mental illness in the press. Cambridge University Press. 10.1017/9781108991278
    https://doi.org/10.1017/9781108991278 [Google Scholar]
  31. Scott, M.
    (2004) WordSmith Tools (Version 4.0) [Computer software]. Oxford University Press. https://lexically.net/wordsmith/downloads/
    [Google Scholar]
  32. (2022) WordSmith Tools (Version 8.0.0.205) [Computer software]. Lexical Analysis Software. https://lexically.net/wordsmith/downloads/
    [Google Scholar]
  33. Sinclair, J. M.
    (2004) Trust the text: Language, corpus and discourse. Routledge. 10.4324/9780203594070
    https://doi.org/10.4324/9780203594070 [Google Scholar]
  34. (2005) Corpus and text — Basic principles. InM. Wynne (Ed.), Developing linguistic corpora. A guide to good practice (pp.–). Oxbow. https://users.ox.ac.uk/~martinw/dlc/chapter1.htm
    [Google Scholar]
  35. Sketch Engine
    Sketch Engine. (n.d.). Deduplication. Lexical Computing. https://www.sketchengine.eu/my_keywords/deduplication/
    [Google Scholar]
  36. Taboada, M.
    (2024) Reported speech and gender in the news: Who is quoted, how are they quoted, and why it matters. Discourse & Communication, (), –. 10.1177/17504813241281713
    https://doi.org/10.1177/17504813241281713 [Google Scholar]
  37. Vanichkina, D., & Bednarek, M.
    (2022) Australian Obesity Corpus manual. Sydney Informatics Hub, The University of Sydney. https://osf.io/h6n82
    [Google Scholar]
  38. Xia, Y., Huan, C., & García Marrugo, A.
    (2025) Fair or biased? A corpus-based study of Australia’s early COVID-19 media representation of China. Social Semiotics, (), –. 10.1080/10350330.2024.2341394
    https://doi.org/10.1080/10350330.2024.2341394 [Google Scholar]
/content/journals/10.1075/ijcl.24197.bed
Loading
/content/journals/10.1075/ijcl.24197.bed
Loading

Data & Media loading...

  • Article Type: Research Article
Keywords: newspaper corpora ; deduplication ; document similarity
This is a required field
Please enter a valid email address
Approval was successful
Invalid data
An Error Occurred
Approval was partially successful, following selected items could not be processed due to error