Lexical diversity and translation literalness in machine translation: a comparative study of pretrained, fine-tuned, and few-shot approaches

Abstract This paper presents a comparative study of neural machine translation (MT) models and large language models, emphasizing translation literalness, lexical diversity, and translation quality. We evaluate open-source neural machine translation (NMT) models, LLaMa 3.1 Instruct, and GPT-4o across nine language pairs in pretrained, fine-tuned, and few-shot prompting settings. This comprehensive analysis underscores the differential strengths of each model, contributing valuable insights into their respective utility in MT. Our findings in the literary and news domains show that fine-tuning reduces translation literalness, allowing freer translations across models, most strongly for LLaMa, which produces the least literal translations across domains. The effect of fine-tuning on lexical diversity is model-dependent: it decreases consistently for LLaMa but remains stable or is improved for NMT. Few-shot prompting, on the other hand, enhances GPT-4o’s lexical diversity across domains and significantly boosts its translation quality in the literary domain, yielding the highest overall COMET scores. Furthermore, our correlation analysis reveals a distinct domain effect: literary translation evaluation rewards diversity, whereas news translation evaluation favors literalness. Overall, the results highlight complementary strengths: few-shot GPT-4o for high-quality, lexically rich translations with low error rates, and fine-tuned LLaMa for freer, less literal translations.

Authors

Institutions

Publication Details

Journal
Natural language processing.
Published
2026-09-30
DOI
https://doi.org/10.1017/nlp.2026.10041
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Lexical diversity and translation literalness in machine translation: a comparative study of pretrained, fine-tuned, and few-shot approaches

Tunga Güngör, Zeynep Yi̇rmi̇beşoğlu
Natural language processing.
Natural Language Processing Techniques
article

Lexical diversity and translation literalness in machine translation: a comparative study of pretrained, fine-tuned, and few-shot approaches

Tunga Güngör, Zeynep Yi̇rmi̇beşoğlu
article en

Abstract

Abstract This paper presents a comparative study of neural machine translation (MT) models and large language models, emphasizing translation literalness, lexical diversity, and translation quality. We evaluate open-source neural machine translation (NMT) models, LLaMa 3.1 Instruct, and GPT-4o across nine language pairs in pretrained, fine-tuned, and few-shot prompting settings. This comprehensive analysis underscores the differential strengths of each model, contributing valuable insights into their respective utility in MT. Our findings in the literary and news domains show that fine-tuning reduces translation literalness, allowing freer translations across models, most strongly for LLaMa, which produces the least literal translations across domains. The effect of fine-tuning on lexical diversity is model-dependent: it decreases consistently for LLaMa but remains stable or is improved for NMT. Few-shot prompting, on the other hand, enhances GPT-4o’s lexical diversity across domains and significantly boosts its translation quality in the literary domain, yielding the highest overall COMET scores. Furthermore, our correlation analysis reveals a distinct domain effect: literary translation evaluation rewards diversity, whereas news translation evaluation favors literalness. Overall, the results highlight complementary strengths: few-shot GPT-4o for high-quality, lexically rich translations with low error rates, and fine-tuned LLaMa for freer, less literal translations.

Natural language processing.
Boğaziçi University (TR)
Quality Education
Openalex Percentile: Top 9%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.