Lexical diversity and translation literalness in machine translation: a comparative study of pretrained, fine-tuned, and few-shot approaches
Abstract This paper presents a comparative study of neural machine translation (MT) models and large language models, emphasizing translation literalness, lexical diversity, and translation quality. We evaluate open-source neural machine translation (NMT) models, LLaMa 3.1 Instruct, and GPT-4o across nine language pairs in pretrained, fine-tuned, and few-shot prompting settings. This comprehensive analysis underscores the differential strengths of each model, contributing valuable insights into their respective utility in MT. Our findings in the literary and news domains show that fine-tuning reduces translation literalness, allowing freer translations across models, most strongly for LLaMa, which produces the least literal translations across domains. The effect of fine-tuning on lexical diversity is model-dependent: it decreases consistently for LLaMa but remains stable or is improved for NMT. Few-shot prompting, on the other hand, enhances GPT-4o’s lexical diversity across domains and significantly boosts its translation quality in the literary domain, yielding the highest overall COMET scores. Furthermore, our correlation analysis reveals a distinct domain effect: literary translation evaluation rewards diversity, whereas news translation evaluation favors literalness. Overall, the results highlight complementary strengths: few-shot GPT-4o for high-quality, lexically rich translations with low error rates, and fine-tuned LLaMa for freer, less literal translations.
Authors
- Tunga Güngör (ORCID: https://orcid.org/0000-0001-9448-9422)
- Zeynep Yi̇rmi̇beşoğlu (ORCID: https://orcid.org/0000-0001-5579-6489)
Institutions
- Boğaziçi University (TR)
Publication Details
- Journal
- Natural language processing.
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1017/nlp.2026.10041
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00