Wasserstein-Regularized Low-Rank Adaptation (OTLoRA): Mitigating Representation Drift and Catastrophic Forgetting in Small Language Models
ABSTRACT: The decentralized deployment of Small Language Models (SLMs) in domain-specific downstream tasks requires continuous parameter adaptation while strictly preserving fundamental linguistic representations. Low-Rank Adaptation (LoRA) has emerged as the premier Parameter-Efficient Fine-Tuning (PEFT) paradigm due to its minimal trainable parameter footprint (<1%) and negligible storage overhead. However, empirical investigations demonstrate that conventional LoRA remains severely vulnerable to catastrophic forgetting when adapted across substantial distribution shifts—such as transitioning from broad web corpora to nuanced financial discourse. Existing regularization methods, such as Euclidean parameter penalties (L2) or Elastic Weight Consolidation (EWC), penalize weight displacements in parameter space without considering the non-linear, non-isometric geometry of multi-head attention latent representations. To overcome this fundamental bottleneck, we propose OT-LoRA (Wasserstein-Regularized Low-Rank Adaptation), an end-to-end framework that regularizes parameter adaptation directly in the latent representation space using Entropy-Regularized Optimal Transport. By formulating an auxiliary knowledge-preservation objective via the Sinkhorn-Knopp algorithm, OT-LoRA enforces geometric manifold alignment between canonical frozen embeddings and adapted latent representations over an anchor replay distribution. Extensive benchmarking across the Financial PhraseBank (target domain) and Stanford Sentiment Treebank (SST-2 anchor domain) demonstrates that OT-LoRA achieves an outstanding 91.23% target accuracy while reducing the catastrophic forgetting rate from 32.41% (standard LoRA) down to 8.69% (a 73.19% relative reduction), yielding a notable Backward Transfer (BWT) improvement of +20.00% (from -27.33% to -7.33%). Furthermore, vectorized log-domain Sinkhorn iterations guarantee complete numeric stability with less than 6% training latency overhead and zero inference memory penalty. OT-LoRA establishes a theoretically grounded, computationally lightweight paradigm for robust, knowledge-preserving domain adaptation in critical real-world language processing systems. Keywords: Parameter-Efficient Fine-Tuning (PEFT); Low-Rank Adaptation (LoRA); Optimal Transport; Sinkhorn Divergence; Catastrophic Forgetting; Domain Adaptation; Representation Drift; Small Language Models.
Authors
- Nguyen Minh Sang
- Tran Thanh Hoa
Institutions
- Trường ĐH Nguyễn Tất Thành (VN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-17
- DOI
- https://doi.org/10.5281/zenodo.22804908
- Primary Topic
- Domain Adaptation and Few-Shot Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00