Whose Voice Survives? A Multi-Translation Reference-Space Framework for Evaluating Narrator Differentiation in LLM Literary Translation: A Case Study of as I Lay Dying

Evaluating literary translation by large language models (LLMs) remains challenging because semantically plausible outputs may still alter the linguistic and cognitive features that distinguish individual narrators. We address this problem by developing and validating a multi-translation narrator-voice reference framework for evaluating narrator differentiation in AI-generated Chinese translations of William Faulkner’s As I Lay Dying. We first construct an annotated corpus of 380 semantic–functional units from 36 source sections and three professional Chinese translations. Of these, 266 units are used to establish the narrator-voice reference space and 114 are reserved for held-out evaluation. The annotation framework captures narrator voice across syntax and rhythm, repetition and cohesion, modality and stance, and cognitive and perceptual perspective, together with six pairwise relationships among Darl, Cash, Vardaman, and Anse. Using the annotated data and the resulting reference space, we investigate GPT-5.6 and DeepSeek-V4 under general and narrator-aware prompting conditions, yielding 1368 AI translation outputs. Our results show narrator- and dimension-specific patterns of preservation, compression, amplification, and reorganization. Narrator-aware prompting is associated with fewer high-risk cases in several narrator–dimension combinations, although some outputs exceed the relational range observed in the professional translations. Blinded expert validation yields a screening detection rate of 91.2% and a false-positive rate of 40.7%. These results indicate that the proposed framework can support human review of AI-generated literary translations while retaining source-supported variation beyond the observed human reference space, and highlight the importance of structured annotation and expert validation for evaluating narrator-voice preservation in LLM-based literary translation.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-06
DOI
https://doi.org/10.3390/app16199902
Primary Topic
Translation Studies and Practices
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Whose Voice Survives? A Multi-Translation Reference-Space Framework for Evaluating Narrator Differentiation in LLM Literary Translation: A Case Study of as I Lay Dying

Jianping Wang, Siao Geng
Applied Sciences
Translation Studies and Practices
article

Whose Voice Survives? A Multi-Translation Reference-Space Framework for Evaluating Narrator Differentiation in LLM Literary Translation: A Case Study of as I Lay Dying

Jianping Wang, Siao Geng
article en

Abstract

Evaluating literary translation by large language models (LLMs) remains challenging because semantically plausible outputs may still alter the linguistic and cognitive features that distinguish individual narrators. We address this problem by developing and validating a multi-translation narrator-voice reference framework for evaluating narrator differentiation in AI-generated Chinese translations of William Faulkner’s As I Lay Dying. We first construct an annotated corpus of 380 semantic–functional units from 36 source sections and three professional Chinese translations. Of these, 266 units are used to establish the narrator-voice reference space and 114 are reserved for held-out evaluation. The annotation framework captures narrator voice across syntax and rhythm, repetition and cohesion, modality and stance, and cognitive and perceptual perspective, together with six pairwise relationships among Darl, Cash, Vardaman, and Anse. Using the annotated data and the resulting reference space, we investigate GPT-5.6 and DeepSeek-V4 under general and narrator-aware prompting conditions, yielding 1368 AI translation outputs. Our results show narrator- and dimension-specific patterns of preservation, compression, amplification, and reorganization. Narrator-aware prompting is associated with fewer high-risk cases in several narrator–dimension combinations, although some outputs exceed the relational range observed in the professional translations. Blinded expert validation yields a screening detection rate of 91.2% and a false-positive rate of 40.7%. These results indicate that the proposed framework can support human review of AI-generated literary translations while retaining source-supported variation beyond the observed human reference space, and highlight the importance of structured annotation and expert validation for evaluating narrator-voice preservation in LLM-based literary translation.

Applied SciencesVol. 16(19)
Henan Institute of Science and Technology (CN), Henan Agricultural University (CN)
Openalex Percentile: Top 2%
Translation Studies and Practices
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.