1887
Volume 13, Issue 2
  • ISSN 2949-6861
  • E-ISSN: 2949-6845

Abstract

Recent legislative mandates have expanded language access in government services, yet research related to the integration of machine translation (MT) remains limited. This study evaluates the quality and efficiency of Neural Machine Translation (NMT) systems (DeepL, Google Translate) and Large Language Models (LLMs) (GPT-4) in translating government-based legal documents from English to Spanish. Methodologically, the study involved twenty-seven professional translators who conducted human translation (HT), machine translation post-editing (MTPE), and quality evaluation. Translation quality was measured using an adapted Multidimensional Quality Metrics (MQM) framework, while technical and temporal post-editing efforts were measured via keylogging software. The findings indicate that (a) MTPE significantly reduces translation completion time compared to HT; (b) GPT-4, an LLM, achieves higher overall quality scores than traditional NMT engines, including DeepL and Google Translate; and (c) MTPE and HT perform similarly in overall quality. The study underscores the potential of LLM-based translation technologies, combined with professional human post-editing, as a high-quality and efficient solution to meeting growing language access demands in government contexts. These findings offer critical insights for policymakers and translation professionals in public services.

Available under the CC BY 4.0 license.
Loading

Article metrics loading...

/content/journals/10.1075/dt.25018.rod
2026-07-09
2026-08-16
Loading full text...

Full text loading...

/deliver/fulltext/dt.25018.rod.html?itemId=/content/journals/10.1075/dt.25018.rod&mimeType=html&fmt=ahah

References

  1. Alvarez-Vidal, Sergi, and Antoni Oliver
    2023 “Assessing MT with Measures of PE Effort.” Ampersand111: 100125. 10.1016/j.amper.2023.100125
    https://doi.org/10.1016/j.amper.2023.100125 [Google Scholar]
  2. Bajčić, Martina, and Dejana Golenko
    2024 “Applying Large Language Models in Legal Translation: The State of the Art.” International Journal of Language and Law131: 171–196.
    [Google Scholar]
  3. Bates, Douglas, Martin Maechler, Ben Bolker, Steve Walker, Rune Haubo Christensen, Henrik Singmann,
    2015 “Package ‘lme4’.” R package documentation.
    [Google Scholar]
  4. Bojar, Ondřej, Jiří Mírovský, Kateřina Rysová, and Magdaléna Rysová
    2018 “EvalD Reference-Less Discourse Evaluation for WMT18.” In Proceedings of the Third Conference on Machine Translation: Shared Task Papers, 541–545. Brussels: Association for Computational Linguistics. 10.18653/v1/W18‑6432
    https://doi.org/10.18653/v1/W18-6432 [Google Scholar]
  5. Bowker, Lynne
    2025 “Machine Translation Literacy.” InThe Palgrave Encyclopedia of Computer-Assisted Language Learning, edited byL. McCallum and D. Tafazoli. Cham: Palgrave Macmillan. 10.1007/978‑3‑031‑51447‑0_258‑1
    https://doi.org/10.1007/978-3-031-51447-0_258-1 [Google Scholar]
  6. Briva-Iglesias, Vicent
    2021 “Traducción Humana vs. Traducción Automática: Análisis Contrastivo e Implicaciones para la Aplicación de la Traducción Automática en Traducción Jurídica.” Mutatis Mutandis141: 571–600. 10.17533/udea.mut.v14n2a14
    https://doi.org/10.17533/udea.mut.v14n2a14 [Google Scholar]
  7. 2024Fostering Human-Centered, Augmented Machine Translation: Analysing Interactive Post-Editing. PhD dissertation, Dublin City University, Ireland. https://doras.dcu.ie/30182/
    [Google Scholar]
  8. Briva-Iglesias, Vicent, João Luis Camargo, and Gokhan Dogru
    2024 “Large Language Models ‘Ad Referendum’: How Good Are They at Machine Translation in the Legal Domain?” MonTI. Monografías de Traducción e Interpretación161: 75–107. 10.6035/MonTI.2024.16.02
    https://doi.org/10.6035/MonTI.2024.16.02 [Google Scholar]
  9. Briva-Iglesias, Vicent, Benjamin R. Cowan, and Sharon O’Brien
    2023 “The Impact of Traditional and Interactive Post-Editing on Machine Translation User Experience, Quality, and Productivity.” Translation, Cognition & Behavior61: 58–84. 10.1075/tcb.00077.bri
    https://doi.org/10.1075/tcb.00077.bri [Google Scholar]
  10. Carl, Michael
    2012 “Translog-II: A Program for Recording User Activity Data for Empirical Translation Process Research.” Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC 2012). Available at: www.lrec-conf.org/proceedings/lrec2012/index.html
    [Google Scholar]
  11. Carpuat, Marine, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao, Alvin Grissom II, Marzena Karpinska, Elaine Khoong, William Lewis, André Martins, Mary Nurminen, Douglas Oard, Maja Popović, Michel Simard, and François Yvon
    2025 “An Interdisciplinary Approach to Human-Centered Machine Translation.” Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 22859–22879. https://aclanthology.org/2025.emnlp-main.1164/.
    [Google Scholar]
  12. Castaldo, Antonio, Sheila Castilho, Joss Moorkens, and Johanna Monti
    2025 “Extending CREAMT: Leveraging Large Language Models for Literary Translation Post-Editing.” Proceedings of Machine Translation Summit XX: Volume11, 506–515. https://aclanthology.org/2025.mtsummit-1.40/.
    [Google Scholar]
  13. Castilho, Sheila, and Helena de Medeiros Caseli
    2023 “Tradução Automática.” InEloize R. Marques Seno , eds., Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português. Brasileiras em PLN.
    [Google Scholar]
  14. Chichirau, Malina, Rik van Noord, and Antonio Toral
    2023 “Automatic Discrimination of Human and Neural Machine Translation in Multilingual Scenarios.” InProceedings of the 24th Annual Conference of the European Association for Machine Translation, 217–226. Tampere, Finland: European Association for Machine Translation. https://aclanthology.org/2023.eamt-1.21/
    [Google Scholar]
  15. Cui, Ying, Xiao Liu, and Yuqin Cheng
    2023 “A Comparative Study on the Effort of Human Translation and Post-Editing in Relation to Text Types: An Eye-Tracking and Key-Logging Experiment.” SAGE Open131. 10.1177/21582440231155849
    https://doi.org/10.1177/21582440231155849 [Google Scholar]
  16. De Camillis, Flavia, Egon Stemle, Elena Chiocchetti, and Francesco Fernicola
    2023 “The MT@BZ Corpus: Machine Translation & Legal Language.” In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, 171–180. Tampere: EAMT. https://aclanthology.org/2023.eamt-1.0/
    [Google Scholar]
  17. Drugan, Joanna, Ingemar Strandvik, and Erkka Vuorinen
    2018 “Translation Quality, Quality Management and Agency: Principles and Practice in the European Union Institutions.” InTranslation Quality Assessment: From Principles to Practice, edited byJoss Moorkens, Sheila Castilho, Federico Gaspari, and Stephen Doherty, 39–68. Cham: Springer International Publishing. 10.1007/978‑3‑319‑91241‑7_3
    https://doi.org/10.1007/978-3-319-91241-7_3 [Google Scholar]
  18. ELIS Research
    ELIS Research 2025 European Language Industry Survey (ELIS) 2025. Report. https://elis-survey.org/repository/.
    [Google Scholar]
  19. Freitag, Markus, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey
    2021a “Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation.” Transactions of the Association for Computational Linguistics91: 1460–1474. 10.1162/tacl_a_00437
    https://doi.org/10.1162/tacl_a_00437 [Google Scholar]
  20. Freitag, Markus, Ricardo Rei, Nikita Mathur, Chi-kiu Lo, Craig Stewart, George Foster,
    2021b “Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-Based Human Evaluations on TED and News Domain.” In Proceedings of the Sixth Conference on Machine Translation, 733–774. Association for Computational Linguistics. https://aclanthology.org/2021.wmt-1.73/
    [Google Scholar]
  21. Freitag, Markus, Nitika Mathur, Daniel Deutsch, Chi-Kiu Lo, Eleftherios Avramidis, Ricardo Rei, and Alon Lavie
    2024 “Are LLMs Breaking MT Metrics? Results of the WMT24 Metrics Shared Task.” InProceedings of the Ninth Conference on Machine Translation, 47–81. 10.18653/v1/2024.wmt‑1.2
    https://doi.org/10.18653/v1/2024.wmt-1.2 [Google Scholar]
  22. GALA (Globalization and Localization Association)
    GALA (Globalization and Localization Association) 2025 “GALA Report: Technology, AI and Automation 2025.” https://resources.gala-global.org/business-barometer-report-ai-2025/
    [Google Scholar]
  23. Giampieri, Patrizia
    2023aLegal Machine Translation Explained: MT in Legal Contexts. Newcastle upon Tyne: Cambridge Scholars Publishing.
    [Google Scholar]
  24. 2023b “Is Machine Translation Reliable in the Legal Field? A Corpus-Based Critical Comparative Analysis for Teaching ESP at Tertiary Level.” ESP Today11 (1): 119–137. 10.18485/esptoday.2023.11.1.6
    https://doi.org/10.18485/esptoday.2023.11.1.6 [Google Scholar]
  25. 2025 “Assessing the Quality of AI and MT in Legal Translation.” Altre Modernità: 143–159. 10.54103/2035‑7680/28886
    https://doi.org/10.54103/2035-7680/28886 [Google Scholar]
  26. Hajek, John, Anthony Pym, Yu Hao, Maria Karidakis, Ambrin Hasnain, Anila Hasnain, Juerong Qiu, Ke Hu, and Rachel Macreadie
    2024 “Understanding and Improving Machine Translations for Emergency Communications.” University of Melbourne, Australia. 10.17613/jthe‑m639
    https://doi.org/10.17613/jthe-m639 [Google Scholar]
  27. Jakobsen, Arnt Lykke
    2017 “Translation Process Research.” InThe Handbook of Translation and Cognition, edited byJohn W. Schwieter and Aline Ferreira, 19–49. Hoboken, NJ: Wiley. 10.1002/9781119241485.ch2
    https://doi.org/10.1002/9781119241485.ch2 [Google Scholar]
  28. Kenny, Dorothy
    (ed.) 2022 “Human and Machine Translation.” InMachine Translation for Everyone: Empowering Users in the Age of Artificial Intelligence, 18–23.
    [Google Scholar]
  29. Killman, Jeffrey
    2023 “Machine Translation and Legal Terminology: Data-driven Approaches to Contextual Accuracy.” Handbook of Terminology (Vol.21), edited byŁucja Biel and Hendrik J. Kockaert, 485–510. Amsterdam: Philadelphia.
    [Google Scholar]
  30. Killman, Jeffrey, and Christopher D. Mellinger
    2022 “Technologized Legal Translation and Interpreting: Resource Potential, Availability, and Applications.” Revista de Llengua i Dret781:1–8. https://revistes.eapc.gencat.cat/index.php/rld/issue/view/n78
    [Google Scholar]
  31. Killman, Jeffrey, and Mónica Rodríguez-Castro
    2022 “Post-editng vs. Translating in the Legal Context: Quality and Time Effects from English to Spanish.” Revista de Llengua i Dret781: 56–72. https://revistes.eapc.gencat.cat/index.php/rld/issue/view/n78
    [Google Scholar]
  32. Kincaid, J. Peter, Robert P. Fishburne Jr., Richard L. Rogers and Brad S. Chissom
    1975 Derivation of new readability formulas (Automated Readability Index, Fog Count and Fleisch Reading Ease Formula) for Navy enlisted personnel. Institute for Simulation and Training. https://stars.library.ucf.edu/istlibrary/56
    [Google Scholar]
  33. Krings, Hans P.
    2001Repairing Texts: Empirical Investigations of Machine Translation Post-Editing Processes. Kent, OH: Kent State University Press.
    [Google Scholar]
  34. Kuznetsova, Alexandra, Per B. Brockhoff, and Rune H. B. Christensen
    2017 “lmerTest Package: Tests in Linear Mixed Effects Models.” Journal of Statistical Software821: 1–26. 10.18637/jss.v082.i13
    https://doi.org/10.18637/jss.v082.i13 [Google Scholar]
  35. Lankford, Séamus, Haithem Afli, and Andy Way
    2023 “adaptMLLM: Fine-Tuning Multilingual Language Models on Low-Resource Languages with Integrated LLM Playgrounds.” Information141: 638. 10.3390/info14120638
    https://doi.org/10.3390/info14120638 [Google Scholar]
  36. Läubli, Samuel, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, and Antonio Toral
    2020 “A Set of Recommendations for Assessing Human-Machine Parity in Language Translation.” Journal of Artificial Intelligence Research671: 653–672. 10.1613/jair.1.11371.
    https://doi.org/10.1613/jair.1.11371 [Google Scholar]
  37. Lipski, John M.
    2008 Varieties of Spanish in the United States. Georgetown University Press.
    [Google Scholar]
  38. Liu, Zhongtao, Parker Riley, Daniel Deutsch, Alison Lui, Mengmeng Niu, Apurva Shah, and Markus Freitag
    2024 “Beyond Human-Only: Evaluating Human-Machine Collaboration for Collecting High-Quality Translation Data.” In Proceedings of the Ninth Conference on Machine Translation, 1095–1106. Association for Computational Linguistics.
    [Google Scholar]
  39. Lommel, Arle
    2018 “Metrics for Translation Quality Assessment: A Case for Standardising Error Typologies.” InTranslation Quality Assessment, edited byJoss Moorkens, Sheila Castilho, Federico Gaspari, and Stephen Doherty, 109–127. Cham: Springer.
    [Google Scholar]
  40. Lommel, Arle, Hans Uszkoreit, and Aljoscha Burchardt
    2014 “Multidimensional Quality Metrics (MQM): A Framework for Declaring and Describing Translation Quality Metrics.” Tradumàtica121: 455–463.
    [Google Scholar]
  41. Moorkens, Joss, and Ana Guerberof Arenas
    2024 “Artificial Intelligence, Automation and the Language Industry.” InHandbook of the Language Industry, 71–98.
    [Google Scholar]
  42. Moorkens, Joss, Andy Way, and Séamus Lankford
    2024Automating Translation. Chapter 11: “Sociotechnical Effects of Machine Translation,” 309–330. Routledge.
    [Google Scholar]
  43. Navarro, Ángel, Miquel Domingo, and Francisco Casacuberta
    2023 “Segment-Based Interactive Machine Translation at a Character Level.” In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, 239–248.
    [Google Scholar]
  44. Nimdzi
    Nimdzi 2025 The 2025 Nimdzi 100 Preliminary Ranking. Online report.
    [Google Scholar]
  45. Ortega, Pedro, Tiffany M. Shin, and Gabriela A. Martínez
    2022 “Rethinking the Term ‘Limited English Proficiency’ to Improve Language-Appropriate Healthcare for All.” Journal of Immigrant and Minority Health24(3): 799–805.
    [Google Scholar]
  46. Popović, Maja
    2018 “Error Classification and Analysis for Machine Translation Quality Assessment.” InTranslation Quality Assessment: From Principles to Practice, 129–158. Springer.
    [Google Scholar]
  47. Quinci, Carla, and Gianluca Pontrandolfo
    2023 “Testing Neural Machine Translation Against Different Levels of Specialisation: An Exploratory Investigation Across Legal Genres and Languages.” Trans-kom16(1): 174–209. https://trans-kom.eu/ihv_16_01_2023.html
    [Google Scholar]
  48. R Core Team
    R Core Team 2022 R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna. https://www.R-project.org
    [Google Scholar]
  49. Rivera-Trigueros, Irene
    2022 “Machine Translation Systems and Quality Assessment: A Systematic Review.” Language Resources and Evaluation56(2): 593–619. 10.1007/s10579‑021‑09537‑5
    https://doi.org/10.1007/s10579-021-09537-5 [Google Scholar]
  50. Robinson, Nathaniel, Perez Ogayo, David R. Mortensen, and Graham Neubig
    2023 “ChatGPT MT: Competitive for High- (but Not Low-) Resource Languages.” Proceedings of the 8th Conference on Machine Translation, 392–418. 10.18653/v1/2023.wmt‑1.40
    https://doi.org/10.18653/v1/2023.wmt-1.40 [Google Scholar]
  51. Sarti, Gabriele, Vilém Zouhar, Grzegorz Chrupała, Ana Guerberof-Arenas, Malvina Nissim, and Arianna Bisazza
    2025 “QE4PE: Word-Level Quality Estimation for Human Post-Editing.” Transactions of the Association for Computational Linguistics131: 1410–1435. 10.1162/tacl.a.46
    https://doi.org/10.1162/tacl.a.46 [Google Scholar]
  52. Sizov, Fedor, Cristina España-Bonet, Josef van Genabith, Ruilin Xie, and Khandaker D. Chowdhury
    2024 “Analysing Translation Artifacts: A Comparative Study of LLMs, NMTs, and Human Translations.” InProceedings of the Ninth Conference on Machine Translation, 1183–1199. 10.18653/v1/2024.wmt‑1.116
    https://doi.org/10.18653/v1/2024.wmt-1.116 [Google Scholar]
  53. Snover, Matthew, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul
    2006 “A Study of Translation Edit Rate with Targeted Human Annotation.” InProceedings of the 7th Conference of the Association for Machine Translation in the Americas: Technical Papers, 223–231. https://aclanthology.org/2006.amta-papers.25/
    [Google Scholar]
  54. Terribile, Stefania
    2023 “Is Post-Editing Really Faster than Human Translation?” Translation Spaces13(2): 171–199. 10.1075/ts.22044.ter
    https://doi.org/10.1075/ts.22044.ter [Google Scholar]
  55. Truong, Anh
    2022 “Linguistic Legal Deserts: Addressing Language Access in the United States Legal System for Limited English Proficient Asian Americans and Pacific Islanders.” Boston University Law Review102(4): 1441–1489. https://www.bu.edu/bulawreview/2022/05/26/volume-102-number-4/
    [Google Scholar]
  56. U.S. Census Bureau
    U.S. Census Bureau 2022 Language use in the United States: 2019 (American Community Survey Reports, ACS-50). U.S. Census Bureau. https://www.census.gov/library/publications/2022/acs/acs-50.html
    [Google Scholar]
  57. Vieira, Lucas N., Minako O’Hagan, and Carol O’Sullivan
    2021 “Understanding the Societal Impacts of Machine Translation: A Critical Review of the Literature on Medical and Legal Use Cases.” Information, Communication & Society24(11): 1515–1532. 10.1080/1369118X.2020.1776370
    https://doi.org/10.1080/1369118X.2020.1776370 [Google Scholar]
  58. Vigier-Moreno, Francisco J., and Lorena Pérez-Macías
    2022 “Assessing Neural Machine Translation of Court Documents: A Case Study on the Translation of a Spanish Remand Order into English.” Revista de Llengua i Dret781: 73–91. https://revistes.eapc.gencat.cat/index.php/rld/issue/view/n78
    [Google Scholar]
  59. Wallace, Melissa, and Esther Monzó-Nebot
    2019 “Legal Translation and Interpreting in Public Services: Defining Key Issues, Re-Examining Policies, and Locating the Public in Public Service Interpreting and Translation.” Revista de Llengua i Dret711:1–12. https://revistes.eapc.gencat.cat/index.php/rld/issue/view/n71
    [Google Scholar]
  60. Wang, Haifeng, Hua Wu, Zhaopeng He, Liang Huang, and Kenneth W. Church
    2022 “Progress in Machine Translation.” Engineering18(11): 143–153. https://www.engineering.org.cn/engi/EN/10.1016/j.eng.2021.03.023
    [Google Scholar]
  61. Wei, Johnny, and Robin Jia
    2021 “The Statistical Advantage of Automatic NLG Metrics at the System Level.” InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics, 6840–6854. Association for Computational Linguistics. 10.18653/v1/2021.acl‑long.533
    https://doi.org/10.18653/v1/2021.acl-long.533 [Google Scholar]
  62. Wickham, Hadley, Mara Averick, Jennifer Bryan, Winston Chang, Lucy D’Agostino McGowan, Romain François,
    2019 “Welcome to the Tidyverse.” Journal of Open Source Software4(43): 1686. 10.21105/joss.01686
    https://doi.org/10.21105/joss.01686 [Google Scholar]
  63. Wickham, Hadley, and Jennifer Bryan
    2023 R Packages. 2nd ed. Sepastopol, CA: O’Reilly Media.
    [Google Scholar]
  64. Wiesmann, Eva
    2019 “Machine Translation in the Field of Law: A Study of the Translation of Italian Legal Texts into German.” Comparative Legilinguistics37(1): 117–153. 10.14746/cl.2019.37.4
    https://doi.org/10.14746/cl.2019.37.4 [Google Scholar]
/content/journals/10.1075/dt.25018.rod
Loading
/content/journals/10.1075/dt.25018.rod
Loading

Data & Media loading...

This is a required field
Please enter a valid email address
Approval was successful
Invalid data
An Error Occurred
Approval was partially successful, following selected items could not be processed due to error