Adaptation-Preserving Structured Pruning for Efficient Transfer Learning in Deep Neural Networks
Transfer learning enables pretrained deep networks to adapt to new tasks but typically preserves their original computational structure, leaving much of the inference cost unchanged. Structured pruning can reduce this complexity; however, introducing sparsity during fine-tuning creates a competing optimization problem in which transferred representations must adapt while the network simultaneously determines which structures can be removed. We propose GASP, an adaptation-preserving structured pruning framework that treats task adaptation and structural sparsity learning as a coordinated optimization process, allowing for the sparse structure to emerge progressively as pretrained representations adapt to the target task. Rather than imposing structural decisions independently of this adaptation, GASP first establishes target-task adaptation through a warm-up phase, before progressively introducing sparsity pressure through an equation-driven schedule. In parallel, temperature annealing gradually sharpens input-aware continuous gates toward increasingly differentiated structural responses. Together, these mechanisms regulate the onset and progression of sparsification while structural importance evolves during fine-tuning. The learned gating responses are subsequently aggregated over the validation data and converted into a physically compact architecture through global Otsu-based thresholding and structural reconstruction, avoiding predefined sparsity ratios and iterative prune–retrain cycles. The same optimization principle is applied to channels in CNNs and MLP units in Swin Transformers without additional trainable mask parameters. Experiments with VGG16, EfficientNet-B0, and Swin-Tiny on CIFAR-10, PlantVillage, and PlantDoc show that GASP achieves accuracy improvements of up to 2.22%, with the corresponding configuration reducing parameters and FLOPs by 91.37% and 88.83%, respectively. Raspberry Pi 4 deployment further demonstrates that the learned structural sparsity translates into practical inference efficiency. Overall, the results demonstrate the effectiveness of coordinating representation adaptation and structural sparsification when deriving compact task-adapted models.
Authors
- Mohammed El Ghzaoui (ORCID: https://orcid.org/0000-0003-3416-2246)
- Mouhssine Chahbouni
- Rachid El Alami (ORCID: https://orcid.org/0000-0002-8524-9053)
- Badre Bossoufi (ORCID: https://orcid.org/0000-0001-8126-7804)
- Imane El Manaa
- Khaoula El Manaa (ORCID: https://orcid.org/0009-0007-2810-4871)
- Anouar Chahbouni (ORCID: https://orcid.org/0009-0005-4333-8698)
- Yassine Abouch
Institutions
- Sidi Mohamed Ben Abdellah University (MA)
Publication Details
- Journal
- Computers
- Published
- 2026-10-04
- DOI
- https://doi.org/10.3390/computers15100679
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00