English Translation Quality Estimation And Fine-Grained Error Pattern Recognition Based On Multi-Source Corpus Augmentation And Large Language Models
Translation quality estimation (QE) must support both sentence-level scoring and local error diagnosis under limited, heterogeneous supervision. This study proposes MCA-LLM-QEER, which combines leakage-controlled multi-source corpus augmentation, source-language/crosslingual/ target-language semantic evidence, explicit linguistic features, and structured LLM evidence in a joint prediction framework. On WMT'23-QE, the model achieves an average Spearman correlation of 0.687 and a Pearson correlation of 0.701, with MAE and RMSE of 0.132 and 0.178. On Domain-QE, error-span F1, error-type Macro-F1, and severity Macro-F1 reach 0.759, 0.708, and 0.724, respectively. Cross-domain and source-target mismatch tests further show that the framework improves robustness to domain shift and fluent-butunfaithful translations. The results support a unified reference-free QE design in which global quality prediction and fine-grained error diagnosis are learned from complementary bilingual evidence rather than from target-side fluency alone.
Authors
- Wei Wang (ORCID: https://orcid.org/0000-0003-0718-5995)
- Weifeng Liu (ORCID: https://orcid.org/0009-0004-0106-9306)
Institutions
- Twitter (United States) (US)
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1142/s0218001426400586
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00