Comparing LoRA, DoRA, and OFT for Extremely Low-Resource English–Sango Machine Translation: Adaptation Performance and Out-of-Domain Retention
Extremely low-resource machine translation poses a difficult cross-lingual adaptation problem: improving target-language performance while preserving capabilities acquired through multilingual pretraining. This study compares three mathematically distinct parameter-efficient fine-tuning (PEFT) methods—LoRA, DoRA, and orthogonal fine-tuning (OFT)—for English–Sango translation using NLLB-200-distilled-600M, a deterministically constructed 9000-pair corpus, and three random seeds per method. Following standard domain-adaptation practice, adaptation-test performance and out-of-domain (OOD) change were evaluated separately using a native-speaker-validated adaptation test set and FLORES+ devtest, with corpus chrF as the primary metric and paired bootstrap testing with Holm correction. All nine adapted runs scored below zero-shot NLLB-200 on the complete adaptation test and FLORES+ devtest. However, the adaptation-side estimate was composition-sensitive: the descriptive direction reversed on the 38-item A-only subset, and checkpoint selection used an all-A development set whereas the final test was B-dominated. The complete adaptation-test result is therefore treated as exploratory. By contrast, the negative FLORES+ pattern was consistent across all nine adapted runs and persisted at fixed final checkpoints, making out-of-domain degradation the most stable performance finding. The methods nevertheless differed computationally: OFT used approximately one quarter of the trainable parameters required by LoRA or DoRA, whereas LoRA achieved the shortest wall-clock training time. These results provide a controlled multi-seed comparison of structurally distinct PEFT configurations for an extremely low-resource English–Sango setting and highlight the importance of considering data quality, zero-shot baselines, training variability, and computational efficiency alongside translation performance.
Authors
- Rakhun Kim (ORCID: https://orcid.org/0009-0004-1312-0076)
Institutions
- Hongik University (KR)
Publication Details
- Journal
- Information
- Published
- 2026-10-09
- DOI
- https://doi.org/10.3390/info17101001
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00