Cross-domain traffic forecasting through multimodal fusion and modality-aware learning

The continuous evolution of intelligent transportation systems (ITS) has driven a paradigm shift in traffic forecasting- from classical statistical models to advanced multimodal and cross-domain deep learning architectures. This paper presents a multimodal fusion and modality-aware learning framework evaluated on two real-world heterogeneous datasets: NGSIM (United States) and TiHAN-V2X (India). The framework integrates kinematic, contextual, and interaction modalities through early, intermediate, and late fusion strategies. For vehicle-type classification, early fusion achieves 0.9848 ± 0.0028 accuracy on TiHAN-V2X and 0.9996 ± 0.0001 on NGSIM(mean ± SD, n = 5 seeds). For short-term speed forecasting, TCN + Attention Early Fusion achieves a mean MAE of 6.638 ± 0.016 km/h at the 100 ms horizon and 10.293 ± 0.033 km/h at the 500 ms horizon across five random seeds, outperforming BiLSTM (6.680 ± 0.025 km/h) and TCN + Attention Late Fusion (6.701 ± 0.062 km/h) at all prediction horizons. A cross-dataset generalisation experiment applies NGSIM-trained classification models directly to TiHAN-V2X without retraining, producing physically consistent kinematic profile mappings across all seven TiHAN scenarios. The framework demonstrates strong missing-modality robustness, maintaining stable forecasting performance even when acceleration and lateral position are structurally absent. Late fusion consistently outperforms early fusion under modality dropout, while early fusion achieves the lowest forecasting MAE when modalities are well-aligned, providing complementary design guidelines for real-world ITS deployment. The results ( https://github.com/Jyoti-Yadav-PhD/multimodal-traffic-fusion ) establish that multimodal deep learning with spatial-temporal graph reasoning and attention-based temporal modelling is a viable foundation for next-generation, domain-generalizable traffic forecasting systems.

Authors

Institutions

Publication Details

Journal
Journal of Intelligent & Fuzzy Systems
Published
2026-10-08
DOI
https://doi.org/10.1177/18758967261494065
Primary Topic
Traffic Prediction and Management Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Cross-domain traffic forecasting through multimodal fusion and modality-aware learning

Neeti Kashyap, Shraddha Arora, Jyoti Yadav
Journal of Intelligent & Fuzzy Systems
Traffic Prediction and Management Techniques
article

Cross-domain traffic forecasting through multimodal fusion and modality-aware learning

Neeti Kashyap, Shraddha Arora, Jyoti Yadav
article en

Abstract

The continuous evolution of intelligent transportation systems (ITS) has driven a paradigm shift in traffic forecasting- from classical statistical models to advanced multimodal and cross-domain deep learning architectures. This paper presents a multimodal fusion and modality-aware learning framework evaluated on two real-world heterogeneous datasets: NGSIM (United States) and TiHAN-V2X (India). The framework integrates kinematic, contextual, and interaction modalities through early, intermediate, and late fusion strategies. For vehicle-type classification, early fusion achieves 0.9848 ± 0.0028 accuracy on TiHAN-V2X and 0.9996 ± 0.0001 on NGSIM(mean ± SD, n = 5 seeds). For short-term speed forecasting, TCN + Attention Early Fusion achieves a mean MAE of 6.638 ± 0.016 km/h at the 100 ms horizon and 10.293 ± 0.033 km/h at the 500 ms horizon across five random seeds, outperforming BiLSTM (6.680 ± 0.025 km/h) and TCN + Attention Late Fusion (6.701 ± 0.062 km/h) at all prediction horizons. A cross-dataset generalisation experiment applies NGSIM-trained classification models directly to TiHAN-V2X without retraining, producing physically consistent kinematic profile mappings across all seven TiHAN scenarios. The framework demonstrates strong missing-modality robustness, maintaining stable forecasting performance even when acceleration and lateral position are structurally absent. Late fusion consistently outperforms early fusion under modality dropout, while early fusion achieves the lowest forecasting MAE when modalities are well-aligned, providing complementary design guidelines for real-world ITS deployment. The results ( https://github.com/Jyoti-Yadav-PhD/multimodal-traffic-fusion ) establish that multimodal deep learning with spatial-temporal graph reasoning and attention-based temporal modelling is a viable foundation for next-generation, domain-generalizable traffic forecasting systems.

Journal of Intelligent & Fuzzy Systems
The NorthCap University (IN)
Openalex Percentile: Top 15%
Traffic Prediction and Management Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Cross-domain traffic forecasting through multimodal fusion and modality-aware learning — Neeti Kashyap, Shraddha Arora, et al. · Journal of Intelligent & Fuzzy Systems (2026) | TGRS Research Map | TGRS