DR-SCA: Directional Encoder Blocks and Semantic-Conditioned Coordinate Attention for UAV Semantic Segmentation

Urban unmanned aerial vehicle (UAV) scenes exhibit strong directional structure, and their foreground composition varies substantially across frames. Coordinate Attention encodes directional information through axis-wise pooling, but its gates are produced without an explicit feature-derived global conditioning path. Motivated by these properties, we propose DR-SCA, a compact RepVGG-style encoder–decoder combining Directional RepVGG (D-Rep) encoder blocks and bottleneck Semantic-Conditioned Coordinate Attention (SCA). In the encoder, D-Rep augments the RepVGG block with the horizontal and vertical asymmetric training branches of ACNet, which fold into one deployed 3×3 convolution. SCA applies feature-wise linear modulation (FiLM)-style feature-wise affine modulation to the Coordinate Attention descriptor using a global descriptor from the same bottleneck feature. The contribution lies in the architectural integration and controlled evaluation of encoder-side branch composition and bottleneck feature conditioning. Under a shared 50-epoch protocol with training from scratch over ten seeds, the compact base achieves 59.96% mean intersection over union (mIoU) on the UAVid validation split. Cross-ablation shows that D-Rep and SCA improve the compact base by 0.83 and 1.01 percentage points, respectively. Their combination yields a gain of 1.80 percentage points and a full-model mIoU of 61.75±0.30%. With 2.050M parameters, DR-SCA has the highest mean mIoU among the six from-scratch methods. It exceeds the second-ranked UNetFormer by 5.76 percentage points with 14.8% of its parameters. Baselines with ImageNet-pretrained encoders reduce this margin on the validation split and match or exceed DR-SCA on the official test split. No statistically significant accuracy difference is detected between D-Rep and either the Diverse Branch Block or an ACNet-style block, or between SCA and Coordinate Attention. Under the same controlled training budget, DR-SCA ranks first among the compared methods on UDD6 and LoveDA. Adding SCA to D-Rep without bottleneck attention improves mIoU on UDD6 but not on LoveDA. Dynamic online augmentation preserves the ranking of the three evaluated models on UAVid.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-24
DOI
https://doi.org/10.3390/electronics15194396
Primary Topic
Advanced Neural Network Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

DR-SCA: Directional Encoder Blocks and Semantic-Conditioned Coordinate Attention for UAV Semantic Segmentation

Wangduk Seo, Jihoon Jeong
Electronics
Advanced Neural Network Applications
article

DR-SCA: Directional Encoder Blocks and Semantic-Conditioned Coordinate Attention for UAV Semantic Segmentation

Wangduk Seo, Jihoon Jeong
article en

Abstract

Urban unmanned aerial vehicle (UAV) scenes exhibit strong directional structure, and their foreground composition varies substantially across frames. Coordinate Attention encodes directional information through axis-wise pooling, but its gates are produced without an explicit feature-derived global conditioning path. Motivated by these properties, we propose DR-SCA, a compact RepVGG-style encoder–decoder combining Directional RepVGG (D-Rep) encoder blocks and bottleneck Semantic-Conditioned Coordinate Attention (SCA). In the encoder, D-Rep augments the RepVGG block with the horizontal and vertical asymmetric training branches of ACNet, which fold into one deployed 3×3 convolution. SCA applies feature-wise linear modulation (FiLM)-style feature-wise affine modulation to the Coordinate Attention descriptor using a global descriptor from the same bottleneck feature. The contribution lies in the architectural integration and controlled evaluation of encoder-side branch composition and bottleneck feature conditioning. Under a shared 50-epoch protocol with training from scratch over ten seeds, the compact base achieves 59.96% mean intersection over union (mIoU) on the UAVid validation split. Cross-ablation shows that D-Rep and SCA improve the compact base by 0.83 and 1.01 percentage points, respectively. Their combination yields a gain of 1.80 percentage points and a full-model mIoU of 61.75±0.30%. With 2.050M parameters, DR-SCA has the highest mean mIoU among the six from-scratch methods. It exceeds the second-ranked UNetFormer by 5.76 percentage points with 14.8% of its parameters. Baselines with ImageNet-pretrained encoders reduce this margin on the validation split and match or exceed DR-SCA on the official test split. No statistically significant accuracy difference is detected between D-Rep and either the Diverse Branch Block or an ACNet-style block, or between SCA and Coordinate Attention. Under the same controlled training budget, DR-SCA ranks first among the compared methods on UDD6 and LoveDA. Adding SCA to D-Rep without bottleneck attention improves mIoU on UDD6 but not on LoveDA. Dynamic online augmentation preserves the ranking of the three evaluated models on UAVid.

ElectronicsVol. 15(19)
Kyonggi University (KR)
Sustainable cities and communities
Openalex Percentile: Top 14%
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.