Catastrophic Forgetting in Parameter-Efficient Fine-Tuning of Small Language Models: A Controlled Comparison of LoRA, Bottleneck Adapters, and Full Fine-Tuning under Sequential Task Learning
Abstract Four small language models of 124M–1.1B parameters were trained on the same five-task stream with five configurations: full fine-tuning, LoRA (r=8 and 16), bottleneck adapters, and LoRA r=8 with 2% replay. We evaluate task accuracy, backward transfer, forgetting, forward transfer, and learning plasticity. Across four backbones, the final test mean accuracy is 55.0% for full fine-tuning, 72.0% for adapters, and 67.6% for LoRA r=8, a significant gap motivating choice of method. Update size generally tracks forgetting, with one exception discussed in Section VI-A. Full fine-tuning improves learning by 1.3 average points per task, trading approximately 1 point of plasticity for 17 points of retention. The same pattern occurs within LoRA, where forgetting increases from 17.9 points at r=2 to 30.5 at r=64, while learning accuracy increases by less than four points. Model scale only helps weakly: increasing backbone size by x9 reduces full-fine-tuning forgetting by only 4.6 points. Under LoRA, rehearsing 2% of previous data reduces forgetting to 8.4 points, at almost no model cost. For sub-2B models trained sequentially, we recommend either bottleneck adapters or low-rank LoRA to avoid feed-forward updates and retain a small replay buffer when previous data are available.
Authors
- Mihir Jadhav
- Jatin Nayak
- Vipul Gavade
- Vedant Rajendra Vare
Institutions
- Vivekanand Education Society's College of Arts, Science and Commerce (Autonomous) (IN)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.22725099
- Primary Topic
- Domain Adaptation and Few-Shot Learning
- Type
- article
- Field-Weighted Citation Impact
- 0.00