XLM-Transformer Models for for Chinese-Japanese Neural Machine Translation under Data Augmentation Technology
Cross-lingual semantic alignment and data scarcity remain two major challenges for low-resource language pairs in neural machine translation (NMT), particularly when developing translation models with crossdomain adaptability. This study proposes a Chinese–Japanese NMT framework that integrates the pre-trained cross-lingual model Cross-lingual Language Model–RoBERTa (XLM-R) with a Dynamic Weighted Iterative Back-Translation for Data Augmentation (DWIB-DA) strategy. The key innovation lies in enhancing crosslingual semantic representations through multilingual shared vocabularies and large-scale pretraining, while constructing high-quality pseudo-parallel corpora through dynamic weight adjustment and adversarial training mechanisms. Specifically, XLM-R achieves bilingual semantic space alignment by leveraging multilingual masked language modeling and shared lexical representations. The proposed DWIB-DA strategy dynamically calculates domain relevance and semantic similarity weights to select and generate high-quality pseudo-parallel data. Furthermore, adversarial training is incorporated to enhance the domain adaptability and robustness of the augmented corpus. By combining cross-lingual pretraining with dynamic data augmentation, the proposed framework effectively addresses the challenges of limited training resources and domain variation in Chinese–Japanese NMT. Experimental results demonstrate that the proposed model achieves a BLEU score of 38.9 on general-domain test sets, outperforming the baseline XLM-Transformer by 3.3 points. Under low-resource conditions with only 20% of the original training data, the model maintains a BLEU score of 31.2, indicating strong translation capability with limited supervision. In a domain-specific popular science translation task, the proposed model improves the BLEU score by 4.1 points to 33.0, demonstrating enhanced domain adaptability in specialized translation scenarios. Human evaluation further confirms the superiority of the proposed model in terms of fluency (4.3), semantic accuracy (4.2), and terminological consistency (4.0) compared with the baseline model. These findings provide theoretical insights into the synergistic effects of cross-lingual pretraining and dynamic data augmentation, while offering practical technical references for the development of translation systems targeting similar low-resource language pairs.
Authors
- Yuping Chen (ORCID: https://orcid.org/0009-0002-8115-4900)
Publication Details
- Journal
- International Journal of Pattern Recognition and Artificial Intelligence
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1142/s0218001426400562
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00