DR-SCA: Directional Encoder Blocks and Semantic-Conditioned Coordinate Attention for UAV Semantic Segmentation
Urban unmanned aerial vehicle (UAV) scenes exhibit strong directional structure, and their foreground composition varies substantially across frames. Coordinate Attention encodes directional information through axis-wise pooling, but its gates are produced without an explicit feature-derived global conditioning path. Motivated by these properties, we propose DR-SCA, a compact RepVGG-style encoder–decoder combining Directional RepVGG (D-Rep) encoder blocks and bottleneck Semantic-Conditioned Coordinate Attention (SCA). In the encoder, D-Rep augments the RepVGG block with the horizontal and vertical asymmetric training branches of ACNet, which fold into one deployed 3×3 convolution. SCA applies feature-wise linear modulation (FiLM)-style feature-wise affine modulation to the Coordinate Attention descriptor using a global descriptor from the same bottleneck feature. The contribution lies in the architectural integration and controlled evaluation of encoder-side branch composition and bottleneck feature conditioning. Under a shared 50-epoch protocol with training from scratch over ten seeds, the compact base achieves 59.96% mean intersection over union (mIoU) on the UAVid validation split. Cross-ablation shows that D-Rep and SCA improve the compact base by 0.83 and 1.01 percentage points, respectively. Their combination yields a gain of 1.80 percentage points and a full-model mIoU of 61.75±0.30%. With 2.050M parameters, DR-SCA has the highest mean mIoU among the six from-scratch methods. It exceeds the second-ranked UNetFormer by 5.76 percentage points with 14.8% of its parameters. Baselines with ImageNet-pretrained encoders reduce this margin on the validation split and match or exceed DR-SCA on the official test split. No statistically significant accuracy difference is detected between D-Rep and either the Diverse Branch Block or an ACNet-style block, or between SCA and Coordinate Attention. Under the same controlled training budget, DR-SCA ranks first among the compared methods on UDD6 and LoveDA. Adding SCA to D-Rep without bottleneck attention improves mIoU on UDD6 but not on LoveDA. Dynamic online augmentation preserves the ranking of the three evaluated models on UAVid.
Authors
- Wangduk Seo (ORCID: https://orcid.org/0000-0003-4806-1614)
- Jihoon Jeong
Institutions
- Kyonggi University (KR)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/electronics15194396
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00