Comparing LoRA, DoRA, and OFT for Extremely Low-Resource English–Sango Machine Translation: Adaptation Performance and Out-of-Domain Retention

Extremely low-resource machine translation poses a difficult cross-lingual adaptation problem: improving target-language performance while preserving capabilities acquired through multilingual pretraining. This study compares three mathematically distinct parameter-efficient fine-tuning (PEFT) methods—LoRA, DoRA, and orthogonal fine-tuning (OFT)—for English–Sango translation using NLLB-200-distilled-600M, a deterministically constructed 9000-pair corpus, and three random seeds per method. Following standard domain-adaptation practice, adaptation-test performance and out-of-domain (OOD) change were evaluated separately using a native-speaker-validated adaptation test set and FLORES+ devtest, with corpus chrF as the primary metric and paired bootstrap testing with Holm correction. All nine adapted runs scored below zero-shot NLLB-200 on the complete adaptation test and FLORES+ devtest. However, the adaptation-side estimate was composition-sensitive: the descriptive direction reversed on the 38-item A-only subset, and checkpoint selection used an all-A development set whereas the final test was B-dominated. The complete adaptation-test result is therefore treated as exploratory. By contrast, the negative FLORES+ pattern was consistent across all nine adapted runs and persisted at fixed final checkpoints, making out-of-domain degradation the most stable performance finding. The methods nevertheless differed computationally: OFT used approximately one quarter of the trainable parameters required by LoRA or DoRA, whereas LoRA achieved the shortest wall-clock training time. These results provide a controlled multi-seed comparison of structurally distinct PEFT configurations for an extremely low-resource English–Sango setting and highlight the importance of considering data quality, zero-shot baselines, training variability, and computational efficiency alongside translation performance.

Authors

Institutions

Publication Details

Journal
Information
Published
2026-10-09
DOI
https://doi.org/10.3390/info17101001
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparing LoRA, DoRA, and OFT for Extremely Low-Resource English–Sango Machine Translation: Adaptation Performance and Out-of-Domain Retention

Rakhun Kim
Information
Natural Language Processing Techniques
article

Comparing LoRA, DoRA, and OFT for Extremely Low-Resource English–Sango Machine Translation: Adaptation Performance and Out-of-Domain Retention

Rakhun Kim
article en

Abstract

Extremely low-resource machine translation poses a difficult cross-lingual adaptation problem: improving target-language performance while preserving capabilities acquired through multilingual pretraining. This study compares three mathematically distinct parameter-efficient fine-tuning (PEFT) methods—LoRA, DoRA, and orthogonal fine-tuning (OFT)—for English–Sango translation using NLLB-200-distilled-600M, a deterministically constructed 9000-pair corpus, and three random seeds per method. Following standard domain-adaptation practice, adaptation-test performance and out-of-domain (OOD) change were evaluated separately using a native-speaker-validated adaptation test set and FLORES+ devtest, with corpus chrF as the primary metric and paired bootstrap testing with Holm correction. All nine adapted runs scored below zero-shot NLLB-200 on the complete adaptation test and FLORES+ devtest. However, the adaptation-side estimate was composition-sensitive: the descriptive direction reversed on the 38-item A-only subset, and checkpoint selection used an all-A development set whereas the final test was B-dominated. The complete adaptation-test result is therefore treated as exploratory. By contrast, the negative FLORES+ pattern was consistent across all nine adapted runs and persisted at fixed final checkpoints, making out-of-domain degradation the most stable performance finding. The methods nevertheless differed computationally: OFT used approximately one quarter of the trainable parameters required by LoRA or DoRA, whereas LoRA achieved the shortest wall-clock training time. These results provide a controlled multi-seed comparison of structurally distinct PEFT configurations for an extremely low-resource English–Sango setting and highlight the importance of considering data quality, zero-shot baselines, training variability, and computational efficiency alongside translation performance.

InformationVol. 17(10)
Hongik University (KR)
Openalex Percentile: Top 12%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Comparing LoRA, DoRA, and OFT for Extremely Low-Resource English–Sango Machine Translation: Adaptation Performance and Out-of-Domain Retention — Rakhun Kim · Information (2026) | TGRS Research Map | TGRS