VLM-Mapper: Multimodal vision-language models for prior guided lane-level HD map construction

The automated construction of High-Definition (HD) maps from remote sensing data is essential for modern intelligent transportation systems and spatial data infrastructure. While imagery provides a scalable solution for lane-level HD map construction, traditional discriminative models often struggle in complex scenarios because of their reliance on local visual features and limited use of external geographic context. To address these limitations, this paper introduces VLM-Mapper, a framework that reframes lane-level HD map construction as a multimodal sequence-generation task based on visual–language alignment and prior-conditioned token generation. VLM-Mapper systematically integrates high-resolution aerial imagery with tokenized geographic prior prompts, including standard map metadata and boundary entry points, to guide HD map construction. Furthermore, to mitigate the spatial inconsistencies and count errors of autoregressive decoding, we implement a two-stage post-training pipeline. This strategy establishes spatial alignment via Supervised Fine-Tuning (SFT), followed by a Reinforcement Learning (RL) phase using Group Relative Policy Optimization (GRPO) with a custom geometry-aware and numerical reward function. Experimental results on the large-scale OpenSatMap dataset show that VLM-Mapper achieves state-of-the-art performance and improves the robustness and topological consistency of HD map construction across large-scale aerial imagery.

Authors

Institutions

Publication Details

Journal
International Journal of Applied Earth Observation and Geoinformation
Published
2026-10-07
DOI
https://doi.org/10.1016/j.jag.2026.105632
Primary Topic
Automated Road and Building Extraction
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

VLM-Mapper: Multimodal vision-language models for prior guided lane-level HD map construction

Pengwei Dong, Haofeng Xie, Yibing Xiong, Huiwei Jiang et al.
International Journal of Applied Earth Observation and Geoinformation
Automated Road and Building Extraction
article

VLM-Mapper: Multimodal vision-language models for prior guided lane-level HD map construction

Pengwei Dong, Haofeng Xie, Yibing Xiong, Huiwei Jiang, Yandi Yang, Xiangyun Hu
article en

Abstract

The automated construction of High-Definition (HD) maps from remote sensing data is essential for modern intelligent transportation systems and spatial data infrastructure. While imagery provides a scalable solution for lane-level HD map construction, traditional discriminative models often struggle in complex scenarios because of their reliance on local visual features and limited use of external geographic context. To address these limitations, this paper introduces VLM-Mapper, a framework that reframes lane-level HD map construction as a multimodal sequence-generation task based on visual–language alignment and prior-conditioned token generation. VLM-Mapper systematically integrates high-resolution aerial imagery with tokenized geographic prior prompts, including standard map metadata and boundary entry points, to guide HD map construction. Furthermore, to mitigate the spatial inconsistencies and count errors of autoregressive decoding, we implement a two-stage post-training pipeline. This strategy establishes spatial alignment via Supervised Fine-Tuning (SFT), followed by a Reinforcement Learning (RL) phase using Group Relative Policy Optimization (GRPO) with a custom geometry-aware and numerical reward function. Experimental results on the large-scale OpenSatMap dataset show that VLM-Mapper achieves state-of-the-art performance and improves the robustness and topological consistency of HD map construction across large-scale aerial imagery.

International Journal of Applied Earth Observation and GeoinformationVol. 154
University of Calgary (CA), Wuhan University (CN), State Key Laboratory of Information Engineering in Surveying Mapping and Remote Sensing (CN)
China Association for Science and Technology
Sustainable cities and communities
Openalex Percentile: Top 17%
Automated Road and Building Extraction
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.