Label-Efficient Semi-Supervised Building Semantic Segmentation for High-Resolution UAV Imagery
Accurate building segmentation from high-resolution unmanned aerial vehicle (UAV) imagery is essential for urban mapping, three-dimensional modeling, and related remote sensing applications. However, constructing precise pixel-level annotations for such imagery is labor-intensive, whereas unlabeled imagery from new survey regions can be acquired relatively easily. This study presents a label-efficient semi-supervised framework that jointly utilizes limited labeled source-region imagery and abundant unlabeled target-region imagery without requiring target-region annotations for training. A frozen DINOv2 encoder provides transferable visual representations, and a Dense Prediction Transformer (DPT) decoder reconstructs multi-level features for dense prediction. An exponential moving average (EMA)-based teacher–student strategy incorporates unlabeled target imagery through confidence-controlled pseudo-label learning. Geometry-preserving photometric augmentation and feature-level perturbation improve consistency learning while avoiding artificial disruption of building footprints. Boundary supervision is derived only from labeled masks to enhance building geometry without propagating uncertain pseudo-boundaries. The proposed framework achieved an intersection over union (IoU) of 0.9197 and Boundary IoU of 0.4578 on the Wonju test set, and an IoU of 0.8921 and Boundary IoU of 0.4455 on the Seoul test set. These results demonstrate effective target-region building segmentation under limited annotation conditions.
Authors
- Youkyung Han (ORCID: https://orcid.org/0000-0001-6586-8503)
- Junseo Baek
- Gahyun Lee
Institutions
- Seoul National University of Science and Technology (KR)
Publication Details
- Journal
- Remote Sensing
- Published
- 2026-09-21
- DOI
- https://doi.org/10.3390/rs18183250
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00