Wasserstein-Regularized Low-Rank Adaptation (OTLoRA): Mitigating Representation Drift and Catastrophic Forgetting in Small Language Models

ABSTRACT: The decentralized deployment of Small Language Models (SLMs) in domain-specific downstream tasks requires continuous parameter adaptation while strictly preserving fundamental linguistic representations. Low-Rank Adaptation (LoRA) has emerged as the premier Parameter-Efficient Fine-Tuning (PEFT) paradigm due to its minimal trainable parameter footprint (<1%) and negligible storage overhead. However, empirical investigations demonstrate that conventional LoRA remains severely vulnerable to catastrophic forgetting when adapted across substantial distribution shifts—such as transitioning from broad web corpora to nuanced financial discourse. Existing regularization methods, such as Euclidean parameter penalties (L2) or Elastic Weight Consolidation (EWC), penalize weight displacements in parameter space without considering the non-linear, non-isometric geometry of multi-head attention latent representations. To overcome this fundamental bottleneck, we propose OT-LoRA (Wasserstein-Regularized Low-Rank Adaptation), an end-to-end framework that regularizes parameter adaptation directly in the latent representation space using Entropy-Regularized Optimal Transport. By formulating an auxiliary knowledge-preservation objective via the Sinkhorn-Knopp algorithm, OT-LoRA enforces geometric manifold alignment between canonical frozen embeddings and adapted latent representations over an anchor replay distribution. Extensive benchmarking across the Financial PhraseBank (target domain) and Stanford Sentiment Treebank (SST-2 anchor domain) demonstrates that OT-LoRA achieves an outstanding 91.23% target accuracy while reducing the catastrophic forgetting rate from 32.41% (standard LoRA) down to 8.69% (a 73.19% relative reduction), yielding a notable Backward Transfer (BWT) improvement of +20.00% (from -27.33% to -7.33%). Furthermore, vectorized log-domain Sinkhorn iterations guarantee complete numeric stability with less than 6% training latency overhead and zero inference memory penalty. OT-LoRA establishes a theoretically grounded, computationally lightweight paradigm for robust, knowledge-preserving domain adaptation in critical real-world language processing systems. Keywords: Parameter-Efficient Fine-Tuning (PEFT); Low-Rank Adaptation (LoRA); Optimal Transport; Sinkhorn Divergence; Catastrophic Forgetting; Domain Adaptation; Representation Drift; Small Language Models.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22804908
Primary Topic
Domain Adaptation and Few-Shot Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Wasserstein-Regularized Low-Rank Adaptation (OTLoRA): Mitigating Representation Drift and Catastrophic Forgetting in Small Language Models

Nguyen Minh Sang, Tran Thanh Hoa
Zenodo (CERN European Organization for Nuclear Research)
Domain Adaptation and Few-Shot Learning
article

Wasserstein-Regularized Low-Rank Adaptation (OTLoRA): Mitigating Representation Drift and Catastrophic Forgetting in Small Language Models

Nguyen Minh Sang, Tran Thanh Hoa
article en

Abstract

ABSTRACT: The decentralized deployment of Small Language Models (SLMs) in domain-specific downstream tasks requires continuous parameter adaptation while strictly preserving fundamental linguistic representations. Low-Rank Adaptation (LoRA) has emerged as the premier Parameter-Efficient Fine-Tuning (PEFT) paradigm due to its minimal trainable parameter footprint (<1%) and negligible storage overhead. However, empirical investigations demonstrate that conventional LoRA remains severely vulnerable to catastrophic forgetting when adapted across substantial distribution shifts—such as transitioning from broad web corpora to nuanced financial discourse. Existing regularization methods, such as Euclidean parameter penalties (L2) or Elastic Weight Consolidation (EWC), penalize weight displacements in parameter space without considering the non-linear, non-isometric geometry of multi-head attention latent representations. To overcome this fundamental bottleneck, we propose OT-LoRA (Wasserstein-Regularized Low-Rank Adaptation), an end-to-end framework that regularizes parameter adaptation directly in the latent representation space using Entropy-Regularized Optimal Transport. By formulating an auxiliary knowledge-preservation objective via the Sinkhorn-Knopp algorithm, OT-LoRA enforces geometric manifold alignment between canonical frozen embeddings and adapted latent representations over an anchor replay distribution. Extensive benchmarking across the Financial PhraseBank (target domain) and Stanford Sentiment Treebank (SST-2 anchor domain) demonstrates that OT-LoRA achieves an outstanding 91.23% target accuracy while reducing the catastrophic forgetting rate from 32.41% (standard LoRA) down to 8.69% (a 73.19% relative reduction), yielding a notable Backward Transfer (BWT) improvement of +20.00% (from -27.33% to -7.33%). Furthermore, vectorized log-domain Sinkhorn iterations guarantee complete numeric stability with less than 6% training latency overhead and zero inference memory penalty. OT-LoRA establishes a theoretically grounded, computationally lightweight paradigm for robust, knowledge-preserving domain adaptation in critical real-world language processing systems. Keywords: Parameter-Efficient Fine-Tuning (PEFT); Low-Rank Adaptation (LoRA); Optimal Transport; Sinkhorn Divergence; Catastrophic Forgetting; Domain Adaptation; Representation Drift; Small Language Models.

Zenodo (CERN European Organization for Nuclear Research)
Trường ĐH Nguyễn Tất Thành (VN)
Peace, Justice and strong institutions
Openalex Percentile: Top 8%
Domain Adaptation and Few-Shot Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.