Asymmetric CNN–ViT Residual Fusion for Industrial Surface Defect Detection
Industrial surface inspection must localize defects when annotated data are limited. We present Asymmetric CNN–ViT Residual Fusion (ACVRF), in which a convolutional neural network (CNN) generates proposals and primary region-of-interest features, while a Vision Transformer (ViT) supplies a residual with a bounded mixing coefficient on the same regions. Frozen expert losses determine training-only sample weights. Across three fusion-stage seeds with fixed experts, ACVRF obtained 0.78442 ± 0.00222 mean average precision at an intersection-over-union threshold of 0.50 (mAP50) and 0.37404 ± 0.00089 mAP50:95 on NEU-DET-pro, compared with 0.75955 and 0.35440 for CNN. These results use validation-selected checkpoints. Matched controls support a contribution from the trained ViT/SFP branch, while ACVRF exceeded the strongest tested concatenation baseline by only 0.00458 mAP50. On a supplementary GC10 split with separate validation and test subsets and three independently trained expert pairs, mean test mAP50 was 0.39751 for ACVRF and 0.36804 for CNN. This historically used data pool does not provide new external validation, and the GC10 evidence for a Transformer-specific contribution remains limited. Teacher weighting slightly increased mean mAP50 but reduced mean mAP50:95 on both protocols. Model-only inference required 43.787 ms per image versus 17.936 ms for CNN. These results support CNN-anchored residual fusion within the evaluated protocols, with modest gains over the strongest tested fusion control and a substantial computational cost.
Authors
- Kun Zou
- Yanan Zhang
- Yong Liu
- Yu Zhong
- Chao Ding
Institutions
- Hubei University of Automotive Technology (CN)
Publication Details
- Journal
- Eng—Advances in Engineering
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/eng7090492
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00