Assessment of Convolutional Neural Network and Vision Transformer Architectures for Structural Surface Classification
We evaluate ten convolutional neural networks (CNNs) and ten transformer/attention architectures for nine-category surface classification using 41,756 balanced StructDamage images. A frozen ImageNet-pretrained ResNet18 encoder and learned head use 33,352 training, 4202 validation and 4202 test images grouped by inferred augmentation family. Test accuracy is 98.88%, macro-F1 is 0.9881, and the 95% component-bootstrap accuracy interval is 98.42–99.26%. Content screening retains 4102 test images and 98.85% accuracy. Nevertheless, cross-partition exact matches and candidate source-prefix overlap prevent treating these results as independent-site validation. The separate random-initialization screen uses 1800 training images, 900 validation images, the common test set, 64 × 64 inputs and five epochs. DenseNet121 leads with 88.34% accuracy; XCiT-Tiny leads the transformer/attention group with 84.75% accuracy. Eighteen models select the final epoch, so the comparison does not establish convergence-equivalent superiority. With 5400 training images, all twenty accuracies increase and the between-budget rank correlation is 0.850. Component-aware comparisons and wall/deck error analysis qualify the ranking. The findings concern surface-category recognition, not defect localization, damage severity or structural safety.
Authors
- Bonginkosi Allen Thango (ORCID: https://orcid.org/0000-0003-3635-0988)
- Sipho G. Thango
Institutions
- University of Johannesburg (ZA)
- University of KwaZulu-Natal (ZA)
Publication Details
- Journal
- Journal of Imaging
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/jimaging12100487
- Primary Topic
- Infrastructure Maintenance and Monitoring
- Type
- article
- Field-Weighted Citation Impact
- 0.00