Continuous Cross-Modal Vector Field Learning for Hyperspectral and LiDAR Land-Cover Classification

Multimodal remote sensing classification plays an important role in applications such as urban mapping and environmental monitoring. However, heterogeneous feature distributions across sensing modalities can result in feature misalignment and semantic inconsistency, limiting the effectiveness of multimodal fusion. Existing multimodal fusion approaches, including convolutional, graph-based, attention-based, contrastive, transport-oriented, and sequence-based methods, can exploit complementary information but may not adequately align modality-specific feature distributions. This study proposes FlowErs, a two-stage multimodal classification framework that addresses this challenge through multimodal flow matching. The working hypothesis is that reducing paired feature-distribution discrepancies before feature fusion improves land-cover classification under fixed benchmark protocols. In the first stage, modality-specific MetaFormer encoders extract hierarchical representations from hyperspectral imaging (HSI) and light detection and ranging (LiDAR) data. A multimodal flow matching module then learns time-conditioned velocity fields to transport paired feature distributions toward a shared latent representation using a bidirectional mean-squared-error objective. In the second stage, the pretrained alignment module is reused to align multimodal representations before a lightweight fusion head performs pixel-wise classification. Evaluation on three public benchmark datasets demonstrated a maximum overall accuracy (OA) of 97.56% under fixed benchmark splits. Performance varied across individual land-cover classes, and the study discusses class-specific limitations and the influence of class imbalance on aggregate metrics. The current evaluation is limited to fixed-split experiments without repeated-run statistical analysis or computational profiling. Future work will investigate statistical robustness, computational efficiency, and cross-region generalization.

Authors

Institutions

Publication Details

Journal
Journal of Visualized Experiments
Published
2026-09-08
DOI
https://doi.org/10.3791/72115
Primary Topic
Remote-Sensing Image Classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Continuous Cross-Modal Vector Field Learning for Hyperspectral and LiDAR Land-Cover Classification

Zhiling Suo, Xiangbing Zhu, Bingxin Tian
Journal of Visualized Experiments
Remote-Sensing Image Classification
article

Continuous Cross-Modal Vector Field Learning for Hyperspectral and LiDAR Land-Cover Classification

Zhiling Suo, Xiangbing Zhu, Bingxin Tian
article en

Abstract

Multimodal remote sensing classification plays an important role in applications such as urban mapping and environmental monitoring. However, heterogeneous feature distributions across sensing modalities can result in feature misalignment and semantic inconsistency, limiting the effectiveness of multimodal fusion. Existing multimodal fusion approaches, including convolutional, graph-based, attention-based, contrastive, transport-oriented, and sequence-based methods, can exploit complementary information but may not adequately align modality-specific feature distributions. This study proposes FlowErs, a two-stage multimodal classification framework that addresses this challenge through multimodal flow matching. The working hypothesis is that reducing paired feature-distribution discrepancies before feature fusion improves land-cover classification under fixed benchmark protocols. In the first stage, modality-specific MetaFormer encoders extract hierarchical representations from hyperspectral imaging (HSI) and light detection and ranging (LiDAR) data. A multimodal flow matching module then learns time-conditioned velocity fields to transport paired feature distributions toward a shared latent representation using a bidirectional mean-squared-error objective. In the second stage, the pretrained alignment module is reused to align multimodal representations before a lightweight fusion head performs pixel-wise classification. Evaluation on three public benchmark datasets demonstrated a maximum overall accuracy (OA) of 97.56% under fixed benchmark splits. Performance varied across individual land-cover classes, and the study discusses class-specific limitations and the influence of class imbalance on aggregate metrics. The current evaluation is limited to fixed-split experiments without repeated-run statistical analysis or computational profiling. Future work will investigate statistical robustness, computational efficiency, and cross-region generalization.

Journal of Visualized Experiments(235)
Guilin University of Technology (CN), Shaanxi University of Science and Technology (CN)
Life in Land
Openalex Percentile: Top 13%
Remote-Sensing Image Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.