Automatic Generation of Deep Learning Models Based on Heterogeneous Pre-Trained Model Stitching
Deep learning models have achieved remarkable progress in computer vision tasks, but constructing efficient models still requires substantial expert experience and computational resources. Pre-trained model reuse provides a practical way to reduce model generation cost; however, existing methods still face challenges in feature alignment, structural integration, and resource-constrained search when stitching heterogeneous architectures such as convolutional neural networks (CNNs) and Vision Transformers (ViTs). To address these issues, this paper proposes Multi-Pretrained Model Stitching for Automatic Generation (MPMS-AG), an automatic generation framework for deep learning models based on heterogeneous pre-trained model stitching. MPMS-AG decomposes heterogeneous pre-trained models into reusable neural blocks and formulates model stitching as a sequential decision-making problem. Specifically, it uses hierarchical feature extraction and Radial Basis Function Centered Kernel Alignment (RBF-CKA) to quantify functional similarity between heterogeneous blocks, introduces adaptive block partitioning and hybrid clustering to reduce the search space, and adopts a Generalized Advantage Estimation (GAE)-based Actor–Critic strategy with a Single-Shot Network Pruning (SNIP)-based zero-shot proxy reward to search for stitching paths under parameter and floating point operations (FLOPs). Under the current CIFAR-10 experimental protocol, MPMS-AG generates trainable hybrid architectures with competitive classification performance across homogeneous stitching, heterogeneous cross-architecture stitching, and lightweight model stitching scenarios. The ablation results provide descriptive evidence that hybrid clustering and reinforcement learning-based search contribute to the observed performance. These findings support the feasibility of MPMS-AG for automatic generation of customized deep learning models from heterogeneous pre-trained model libraries within the evaluated setting.
Authors
- Di Cui (ORCID: https://orcid.org/0000-0002-6523-2967)
- Yu Zhao (ORCID: https://orcid.org/0000-0003-0831-3784)
- 彭兰勤
- Shanshan Wu
- Rongping Xie
Institutions
- Xidian University (CN)
- Institute of Electronics (CN)
- Ministry of Industry and Information Technology (CN)
- Nanjing University of Aeronautics and Astronautics (CN)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/app16189215
- Primary Topic
- Advanced Image and Video Retrieval Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Fundamental Research Funds for the Central Universities