1887
image of Data interference
USD
Buy:$35.00 + Taxes

Abstract

Tokenisation is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring the reliability of qualitative approaches. This paper examines how discrepancies in tokenisation affect the representation of language data and the validity of analytical findings. Investigating the challenges posed by emojis and homoglyphs, the study highlights the necessity of pre-processing these elements to maintain corpus fidelity to the source data. The research presents methods for ensuring that digital texts are accurately represented in corpora, thereby supporting reliable linguistic analysis and facilitating the repeatability of linguistic interpretations. The findings emphasise the necessity of a detailed understanding of both linguistic and technical aspects involved in digital textual data to enhance the accuracy of corpus analysis, and have significant implications for both quantitative and qualitative approaches in corpus-based research.

Loading

Article metrics loading...

/content/journals/10.1075/ijcl.24116.dic
2026-06-23
2026-07-17
Loading full text...

Full text loading...

References

  1. Alsulami, A.
    (2019) A sociolinguistic analysis of the use of Arabizi in social media among Saudi Arabians. International Journal of English Linguistics, (), –. 10.5539/ijel.v9n6p257
    https://doi.org/10.5539/ijel.v9n6p257 [Google Scholar]
  2. Andrade, B., Morais, R., & Soares De Lima, E.
    (2024) The personality of visual elements: A framework for the development of visual identity based on brand personality dimensions. The International Journal of Visual Design, (), –. 10.18848/2325‑1581/CGP/v18i01/67‑98
    https://doi.org/10.18848/2325-1581/CGP/v18i01/67-98 [Google Scholar]
  3. Anthony, L.
    (2023) AntConc (4.2.4) [Computer software]. Waseda University. https://www.laurenceanthony.net/software.
    [Google Scholar]
  4. Bertini, F., Rizzo, S. G., & Montesi, D.
    (2019) Can information hiding in social media posts represent a threat?Computer, (), –. 10.1109/MC.2019.2917199
    https://doi.org/10.1109/MC.2019.2917199 [Google Scholar]
  5. Brezina, V.
    (2018) Statistics in corpus linguistics: A practical guide. Cambridge University Press. 10.1017/9781316410899
    https://doi.org/10.1017/9781316410899 [Google Scholar]
  6. Brezina, V., & Platt, W.
    (2024) #LancsBox X (4.0.0) [Computer software]. Lancaster University. lancsbox.lancs.ac.uk.
    [Google Scholar]
  7. Brezina, V., & Timperley, M.
    (2017) How large is the BNC?: A proposal for standardised tokenization and word counting. InProceedings of the 9th international corpus linguistics conference. University of Birmingham. https://www.birmingham.ac.uk/documents/college-artslaw/corpus/conference-archives/2017/general/paper303.pdf.
    [Google Scholar]
  8. Bridge, K., Buck, A., White, S., Wojciakowski, M., Hickey, S., & Radich, Q.
    (2023) Use UTF-8 code pages in Windows apps. Microsoft Learn. https://learn.microsoft.com/en-us/windows/apps/design/globalizing/use-utf8-code-page.
    [Google Scholar]
  9. Calhoun, K., & Fawcett, A.
    (2023) “They edited out her nip nops”: Linguistic innovation as textual censorship avoidance on TikTok. Language@Internet, (), –. 10.14434/li.v21.37371
    https://doi.org/10.14434/li.v21.37371 [Google Scholar]
  10. Chen, R.
    (2019, August30). The sad history of Unicode printf-style format specifiers in Visual C++. The Old New Thing. https://devblogs.microsoft.com/oldnewthing/20190830–00/?p=102823.
    [Google Scholar]
  11. Cooper, P., Surdeanu, M., & Blanco, E.
    (2023) Hiding in plain sight: Tweets with hate speech masked by homoglyphs. InH. Bouamor, J. Pino, & K. Bali (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, –. Association for Computational Linguistics. 10.18653/v1/2023.findings‑emnlp.192
    https://doi.org/10.18653/v1/2023.findings-emnlp.192 [Google Scholar]
  12. Curry, N., Baker, P., & Brookes, G.
    (2024) Generative AI for corpus approaches to discourse studies: A critical evaluation of ChatGPT. Applied Corpus Linguistics, (), . 10.1016/j.acorp.2023.100082
    https://doi.org/10.1016/j.acorp.2023.100082 [Google Scholar]
  13. Daniel, J. S., & Pal, A.
    (2024) Impact of non-Standard Unicode characters on security and comprehension in Large Language Models. arXiv:2405.14490v1. 10.21203/rs.3.rs‑5723808/v1
    https://doi.org/10.21203/rs.3.rs-5723808/v1 [Google Scholar]
  14. Davis, M., & Holbrook, N.
    (Eds.) (2023) UTS #51: Unicode emoji. Unicode Technical Reports. https://www.unicode.org/reports/tr51/.
    [Google Scholar]
  15. Di Cristofaro, M.
    (2023a) Corpus approaches to language in social media. Routledge. 10.4324/9781003225218
    https://doi.org/10.4324/9781003225218 [Google Scholar]
  16. (2023b, July3–6). The hierarchy of web pages: Accounting for contents’ accessibility in keywords analysis [Conference presentation]. The twelfth international corpus linguistics conference 2023, Lancaster, UK.
    [Google Scholar]
  17. Eleta, I., & Golbeck, J.
    (2014) Multilingual use of Twitter: Social networks at the language frontier. Computers in Human Behavior, , –. 10.1016/j.chb.2014.05.005
    https://doi.org/10.1016/j.chb.2014.05.005 [Google Scholar]
  18. Fricke, L., Grosz, P. G., & Scheffler, T.
    (2024) Semantic differences in visually similar face emojis. Language and Cognition, (), –. 10.1017/langcog.2024.12
    https://doi.org/10.1017/langcog.2024.12 [Google Scholar]
  19. Gillings, M., Kohn, T., & Mautner, G.
    (2024) The rise of large language models: Challenges for Critical Discourse Studies. Critical Discourse Studies, (), –. 10.1080/17405904.2024.2373733
    https://doi.org/10.1080/17405904.2024.2373733 [Google Scholar]
  20. Gries, S. T.
    (2016) Quantitative corpus linguistics with R: A practical introduction (2nd ed.). Routledge. 10.4324/9781315746210
    https://doi.org/10.4324/9781315746210 [Google Scholar]
  21. Hafner, C. A.
    (2021) Discourse and computer-mediated communication. InK. Hyland, B. Paltridge, & L. L. C. Wong (Eds.), The Bloomsbury handbook of discourse analysis (2nd ed., pp.–). Bloomsbury Academic. 10.5040/9781350156111.ch‑020
    https://doi.org/10.5040/9781350156111.ch-020 [Google Scholar]
  22. Herring, S. C.
    (Ed.) (1996) Computer-mediated communication: Linguistic, social and cross-cultural perspectives. John Benjamins. 10.1075/pbns.39
    https://doi.org/10.1075/pbns.39 [Google Scholar]
  23. Holtgraves, T., & Robinson, C.
    (2020) Emoji can facilitate recognition of conveyed indirect meaning. PLOS ONE, (), . 10.1371/journal.pone.0232361
    https://doi.org/10.1371/journal.pone.0232361 [Google Scholar]
  24. Horváth, A., van Lit, C., Wagner, C., & Wrisley, D. J.
    (2023) Towards multilingually enabled digital knowledge infrastructures: A qualitative survey analysis. InL. Viola & P. Spence (Eds.), Multilingual digital humanities (pp.–. Routledge. 10.4324/9781003393696‑17
    https://doi.org/10.4324/9781003393696-17 [Google Scholar]
  25. Kilgarriff, A., Baisa, V., Bušta, J., Jakubíček, M., Kovář, V., Michelfeit, J., Rychlý, P., & Suchomel, V.
    (2014) The Sketch Engine: Ten years on. Lexicography, (), –. 10.1007/s40607‑014‑0009‑9
    https://doi.org/10.1007/s40607-014-0009-9 [Google Scholar]
  26. Kim, T.
    (2024) Carpedm20/emoji [Python]. https://github.com/carpedm20/emoji.
    [Google Scholar]
  27. Kohnke, L., Moorhouse, B. L., & Zou, D.
    (2023) ChatGPT for language teaching and learning. RELC Journal, (), –. 10.1177/00336882231162868
    https://doi.org/10.1177/00336882231162868 [Google Scholar]
  28. Konrad, A., Herring, S. C., & Choi, D.
    (2020) Sticker and emoji use in Facebook Messenger: Implications for graphicon change. Journal of Computer-Mediated Communication, (), –. 10.1093/jcmc/zmaa003
    https://doi.org/10.1093/jcmc/zmaa003 [Google Scholar]
  29. Lazarinis, F.
    (2008) Text extraction and web searching in a non-Latin language [Doctoral thesis, University of Sunderland]. sure.sunderland.ac.uk/id/eprint/3326/.
    [Google Scholar]
  30. Logi, L., & Zappavigna, M.
    (2023) A social semiotic perspective on emoji: How emoji and language interact to make meaning in digital messages. New Media & Society, (), –. 10.1177/14614448211032965
    https://doi.org/10.1177/14614448211032965 [Google Scholar]
  31. McCarthy, M. S., & Mothersbaugh, D. L.
    (2002) Effects of typographic factors in advertising‐based persuasion: A general model and initial empirical tests. Psychology & Marketing, (), –. 10.1002/mar.10030
    https://doi.org/10.1002/mar.10030 [Google Scholar]
  32. McEnery, T., & Brezina, V.
    (2022) Fundamental principles of corpus linguistics. Cambridge University Press. 10.1017/9781107110625
    https://doi.org/10.1017/9781107110625 [Google Scholar]
  33. McEnery, T., & Hardie, A.
    (2012) Corpus linguistics: Method, theory and practice. Cambridge University Press.
    [Google Scholar]
  34. McEnery, T., & Xiao, R.
    (2005) Character encoding in corpus construction. InM. Wynne (Ed.), Developing linguistic corpora: A guide to good practice (pp.–. Oxbow Books. https://users.ox.ac.uk/~martinw/dlc/index.htm.
    [Google Scholar]
  35. Miller, H., Thebault-Spieker, J., Chang, S., Johnson, I., Terveen, L., & Hecht, B.
    (2021) “Blissfully happy” or “ready to fight”: Varying interpretations of emoji. Proceedings of the International AAAI Conference on Web and Social Media, (), –. 10.1609/icwsm.v10i1.14757
    https://doi.org/10.1609/icwsm.v10i1.14757 [Google Scholar]
  36. Miller Hillberg, H., Levonian, Z., Kluver, D., Terveen, L., & Hecht, B.
    (2018) What I see is what you don’t get: The effects of (not) seeing emoji rendering differences across platforms. Proceedings of the ACM on Human-Computer Interaction, (), –. 10.1145/3274393
    https://doi.org/10.1145/3274393 [Google Scholar]
  37. Minnich, A., Abu-El-Rub, N., Gokhale, M., Minnich, R., & Mueen, A.
    (2016) ClearView: Data cleaning for online review mining. InR. Kumar, J. Caverlee, & H. Tong (Eds.), 2016 IEEE/ACM international conference on Advances in Social Networks Analysis and Mining (ASONAM), –. Institute of Electrical and Electronics Engineers. 10.1109/ASONAM.2016.7752290
    https://doi.org/10.1109/ASONAM.2016.7752290 [Google Scholar]
  38. Moore, A., & Rayson, P.
    (2022) PyMUSAS: Python Multilingual Ucrel Semantic Analysis System (0.3.0) [Computer software]. https://github.com/ucrel/pymusas.
    [Google Scholar]
  39. Moran, S., & Cysouw, M.
    (2018) The Unicode cookbook for linguists: Managing writing systems using orthography profiles. Zenodo. 10.5281/ZENODO.773250
    https://doi.org/10.5281/ZENODO.773250 [Google Scholar]
  40. Paquot, M., & Gries, S. Th.
    (Eds.) (2020) A Practical handbook of corpus linguistics. Springer. 10.1007/978‑3‑030‑46216‑1
    https://doi.org/10.1007/978-3-030-46216-1 [Google Scholar]
  41. Phillips, A.
    (Ed.) (2021) Character model for the World Wide Web: String matching. World Wide Web Consortium. https://www.w3.org/TR/charmod-norm/.
    [Google Scholar]
  42. Python Software Foundation
    Python Software Foundation (2024) Built-in types. Python Documentation. https://docs.python.org/3/library/stdtypes.html.
    [Google Scholar]
  43. Robertson, A., Magdy, W., & Goldwater, S.
    (2021) Black or White but never neutral: How readers perceive identity from yellow or skin-toned emoji. Proceedings of the ACM on Human-Computer Interaction, (), –. 10.1145/3476091
    https://doi.org/10.1145/3476091 [Google Scholar]
  44. Shoeb, A. A. M., & de Melo, G.
    (2021) Assessing emoji use in modern text processing tools. InC. Zong, F. Xia, W. Li, & R. Navigli (Eds.), Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on Natural Language Processing, –. Association for Computational Linguistics. 10.18653/v1/2021.acl‑long.110
    https://doi.org/10.18653/v1/2021.acl-long.110 [Google Scholar]
  45. Shurick, A. A., & Daniel, J.
    (2020) What’s behind those smiling eyes: Examining emoji sentiment across vendors. InS. Chancellor, K. Garimella, & K. Weller (Eds.), Workshop proceedings of the 14th international AAAI conference on web and social media (Emoji 2020). Association for the Advancement of Artificial Intelligence. 10.36190/2020.04
    https://doi.org/10.36190/2020.04 [Google Scholar]
  46. Spence, P., & Viola, L.
    (2023) Introduction. InL. Viola & P. Spence (Eds.), Multilingual digital humanities (pp.–. Routledge. 10.4324/9781003393696‑1
    https://doi.org/10.4324/9781003393696-1 [Google Scholar]
  47. Spolsky, J.
    (2003, October8). The absolute minimum every software developer absolutely, positively must know about Unicode and character sets (No excuses!). Joel on Software. https://www.joelonsoftware.com/2003/10/08/the-absolute-minimum-every-software-developer-absolutely-positively-must-know-about-unicode-and-character-sets-no-excuses/.
    [Google Scholar]
  48. Storment, J. D.
    (2024) Going ✈ lexicon? The linguistic status of pro-text emojis. Glossa: A Journal of General Linguistics, (). 10.16995/glossa.10449
    https://doi.org/10.16995/glossa.10449 [Google Scholar]
  49. Struppek, L., Hintersdorf, D., Friedrich, F., Brack, M., Schramowski, P., & Kersting, K.
    (2023) Exploiting cultural biases via homoglyphs in text-to-image synthesis. Journal of Artificial Intelligence Research, , –. 10.1613/jair.1.15388
    https://doi.org/10.1613/jair.1.15388 [Google Scholar]
  50. Suzuki, H., Chiba, D., Yoneya, Y., Mori, T., & Goto, S.
    (2019) ShamFinder: An automated framework for detecting IDN homographs. Proceedings of the Internet Measurement Conference, –. 10.1145/3355369.3355587
    https://doi.org/10.1145/3355369.3355587 [Google Scholar]
  51. The TEI Consortium
    The TEI Consortium (2023, November16). TEI P5: Guidelines for electronic text encoding and interchange. TEI Consortium. www.tei-c.org/Guidelines/P5/.
    [Google Scholar]
  52. Uchida, S.
    (2024) Using early LLMs for corpus linguistics: Examining ChatGPT’s potential and limitations. Applied Corpus Linguistics, (), . 10.1016/j.acorp.2024.100089
    https://doi.org/10.1016/j.acorp.2024.100089 [Google Scholar]
  53. Unicode Consortium
    Unicode Consortium (2023) UTS #39: Unicode security mechanisms. Unicode Technical Reports. https://www.unicode.org/reports/tr39/#Confusable_Detection.
    [Google Scholar]
  54. Unicode Consortium
    Unicode Consortium (2024) Supported scripts. Unicode Technical Reports. https://www.unicode.org/standard/supported.html.
    [Google Scholar]
  55. Wang, D., Li, Y., Jiang, J., Ding, Z., Jiang, G., Liang, J., & Yang, D.
    (2024a) Tokenization matters! Degrading Large Language Models through challenging their tokenization. arXiv:2405.17067. 10.48550/arXiv.2405.17067
    https://doi.org/10.48550/arXiv.2405.17067 [Google Scholar]
  56. Wang, Y., Zhang, Y., Zhang, G., Shengyou, H., & Jingsong, Q.
    (2024b) Linguistic properties of emojis: A quantitative exploration of emoji frequency, category, and position on Twitter. Journal of Quantitative Linguistics, (), –. 10.1080/09296174.2024.2347055
    https://doi.org/10.1080/09296174.2024.2347055 [Google Scholar]
  57. Weissman, B., Engelen, J., Baas, E., & Cohn, N.
    (2023) The Lexicon of emoji? Conventionality modulates processing of emoji. Cognitive Science, (), . 10.1111/cogs.13275
    https://doi.org/10.1111/cogs.13275 [Google Scholar]
  58. Whistler, K.
    (2023) UAX #15: Unicode Normalization Forms. Unicode Technical Reports. https://unicode.org/reports/tr15/#Norm_Forms.
    [Google Scholar]
  59. Wiese, H., & Labrenz, A.
    (2021) Emoji as graphic discourse markers. Functional and positional associations in German WhatsApp® messages. InD. Van Olmen & J. Šinkūnienė (Eds.), Pragmatic markers and peripheries (pp.–. John Benjamins. 10.1075/pbns.325.10wie
    https://doi.org/10.1075/pbns.325.10wie [Google Scholar]
  60. Winter, B.
    (2019) Statistics for linguists: An introduction using R. Routledge. 10.4324/9781315165547
    https://doi.org/10.4324/9781315165547 [Google Scholar]
  61. Yu, D., Bondi, M., & Hyland, K.
    (2026) Can GPT-4 learn to analyse moves in research article abstracts?Applied Linguistics, (), –. 10.1093/applin/amae071
    https://doi.org/10.1093/applin/amae071 [Google Scholar]
  62. Zappavigna, M., & Logi, L.
    (2021) Emoji in social media discourse about working from home. Discourse, Context & Media, , . 10.1016/j.dcm.2021.100543
    https://doi.org/10.1016/j.dcm.2021.100543 [Google Scholar]
  63. (2024) Emoji and social media paralanguage. Cambridge University Press. 10.1017/9781009179829
    https://doi.org/10.1017/9781009179829 [Google Scholar]
/content/journals/10.1075/ijcl.24116.dic
Loading
/content/journals/10.1075/ijcl.24116.dic
Loading

Data & Media loading...

  • Article Type: Research Article
Keywords: emojis ; homoglyphs ; UTF-8 ; data fidelity ; tokenisation
This is a required field
Please enter a valid email address
Approval was successful
Invalid data
An Error Occurred
Approval was partially successful, following selected items could not be processed due to error