Preliminary Application Research on Deep Learning Based on a Topic Focusing on the Identification of Modern Cultural and Educational Architectural Styles in Hunan, China
Modern cultural and educational buildings are important material carriers of regional modernization, educational transformation, and Sino-Western cultural exchange. In Hunan, these buildings exhibit complex stylistic features shaped by the interaction of Western architectural languages, Chinese revival forms, modern decorative expressions, and early modern architectural tendencies. This study aims to explore a subject-focused deep learning framework for auxiliary-style recognition of modern educational and cultural buildings in Hunan and to investigate the feasibility of transferring stylistic representations learned from non-Hunan samples to the Hunan target domain under cross-regional conditions. A six-class style classification system with operational morphological criteria was established, and a geographically isolated cross-regional target-domain evaluation framework was adopted. Modern public, educational, and cultural building samples from outside Hunan Province were used for training and validation, while Hunan samples were retained as a fixed target-domain test set. Several representative backbone networks were used as references for model selection, and Swin Transformer was adopted as the unified backbone. The effects of L-Softmax, multi-scale feature fusion, global channel–spatial attention, PFP/PA, GBVS/AGL subject guidance, and Grad-CAM-based background constraints were compared under the unified cross-regional target-domain evaluation protocol. EMA was used only as an auxiliary training stabilization strategy. Model performance was evaluated using image-level accuracy, building-level accuracy based on multi-view majority voting, balanced accuracy, macro-F1, bootstrap confidence intervals, confusion matrices, and representative boundary cases. On the fixed Hunan target-domain test set containing 186 images from 37 building instances, the highest image-level accuracy reached 58.60%, while the highest building-level accuracy reached 64.86%. Considering the strong class imbalance and limited number of building instances in several categories, these results were further interpreted together with balanced metrics and bootstrap confidence intervals rather than accuracy alone. The A1–A8 ablation experiments indicated that the A7 variant incorporating GBVS/AGL achieved a relatively balanced performance between image-level and building-level evaluations under the current target-domain condition. The proposed framework provides an exploratory auxiliary approach for architectural heritage surveys, digital documentation, and style-boundary analysis, rather than a fully generalized architectural style recognition system, while offering methodological reference for future component-level quantitative research on Si-no-Western architectural integration.
Authors
- Yu Yi
- Sheng Song
- Jun Yan
- Jiacheng Liu
- Boyu Pang
- Sumin Li
- Xuchuan Zhou
- Jichi Guo
Institutions
- Changsha University of Science and Technology (CN)
Publication Details
- Journal
- Buildings
- Published
- 2026-09-16
- DOI
- https://doi.org/10.3390/buildings16183678
- Primary Topic
- Advanced Technologies in Various Fields
- Type
- article
- Field-Weighted Citation Impact
- 0.00