STAR: Structure-aware representation learning for efficient and robust UAV-based geo-localization

Abstract Unmanned Aerial Vehicle (UAV)-based geo-localization supports autonomous navigation in GPS-denied or GPS-degraded environments by matching real-time UAV-view images with geo-referenced satellite imagery. A key challenge is to learn a discriminative retrieval space that remains robust to severe cross-view viewpoint discrepancies and temporal shifts. Although visual appearance provides fine-grained discriminative cues, it is sensitive to illumination, seasonal, and perspective variations. In contrast, geographic structures, such as road layouts and building footprints, offer more stable cues, but explicit structural modeling often increases computational cost. To balance robustness and efficiency, we propose STAR, a lightweight STructure-Aware Representation learning framework. STAR uses structural semantics as training-time geometric guidance for visual retrieval manifold learning. Specifically, a structural stream derives structure-aware representations from paired UAV and satellite images and constructs pairwise structural affinities to describe geographic layout relationships. These affinities regularize the discriminative retrieval space at the manifold level, enabling compact visual descriptors to retain structural awareness for robust cross-view matching. We further introduce UniT-1652, a benchmark for evaluating UAV-based geo-localization under temporal shifts. Extensive experiments on University-1652, SUES-200, and UniT-1652 show that STAR achieves state-of-the-art performance across multiple retrieval settings, with improved robustness to temporal and viewpoint variations and fewer parameters than existing methods.

Authors

Publication Details

Journal
Communications in Transportation Research
Published
2026-10-09
DOI
https://doi.org/10.26599/commtr.2026.9640056
Primary Topic
Advanced Image and Video Retrieval Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

STAR: Structure-aware representation learning for efficient and robust UAV-based geo-localization

Yanchen Guan, Zhenning Li, Chengyue Wang, Jiaxun Zhang et al.
Communications in Transportation Research
Advanced Image and Video Retrieval Techniques
article

STAR: Structure-aware representation learning for efficient and robust UAV-based geo-localization

Yanchen Guan, Zhenning Li, Chengyue Wang, Jiaxun Zhang, Xingcheng Liu, Bin Rao, Haicheng Liao
article en

Abstract

Abstract Unmanned Aerial Vehicle (UAV)-based geo-localization supports autonomous navigation in GPS-denied or GPS-degraded environments by matching real-time UAV-view images with geo-referenced satellite imagery. A key challenge is to learn a discriminative retrieval space that remains robust to severe cross-view viewpoint discrepancies and temporal shifts. Although visual appearance provides fine-grained discriminative cues, it is sensitive to illumination, seasonal, and perspective variations. In contrast, geographic structures, such as road layouts and building footprints, offer more stable cues, but explicit structural modeling often increases computational cost. To balance robustness and efficiency, we propose STAR, a lightweight STructure-Aware Representation learning framework. STAR uses structural semantics as training-time geometric guidance for visual retrieval manifold learning. Specifically, a structural stream derives structure-aware representations from paired UAV and satellite images and constructs pairwise structural affinities to describe geographic layout relationships. These affinities regularize the discriminative retrieval space at the manifold level, enabling compact visual descriptors to retain structural awareness for robust cross-view matching. We further introduce UniT-1652, a benchmark for evaluating UAV-based geo-localization under temporal shifts. Extensive experiments on University-1652, SUES-200, and UniT-1652 show that STAR achieves state-of-the-art performance across multiple retrieval settings, with improved robustness to temporal and viewpoint variations and fewer parameters than existing methods.

Communications in Transportation Research
Openalex Percentile: Top 15%
Advanced Image and Video Retrieval Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.