Intrinsic modality disparity augmentation without database via cross-modal routing

Current augmentation of Large Language Models (LLMs) primarily relies on external corpora, which constrains performance based on database availability and leaves models vulnerable to input perturbations. Inspired by the human perceptual system’s ability to leverage mutual augmentation across different modalities, We propose Intrinsic Modality Disparity Augmentation (IMD), a reliability-aware cross-modal re-encoding framework that does not rely on external retrieval databases during inference, while using a labeled calibration set to estimate modality-specific routing centroids. Our approach consists of two core components: first, an on-demand synthesis mechanism that assesses the reliability of the initial query and generates complementary modal representations when text is deemed insufficient; second, a dynamic cross-modal latent space routing module that evaluates and ranks candidates based on their proximity to learned expert centroids. Extensive experiments on diverse reasoning benchmarks demonstrate that IMD significantly outperforms the MLLM’s native text-only inference and standard augmentation methods, delivering substantial gains in both reasoning accuracy and robustness against textual noise. Furthermore, efficiency analysis reveals that IMD maintains a favorable balance between performance and computational overhead, achieving 3x speedups over naive multimodal conversion and remaining highly competitive with traditional text-based augmentation methods.

Authors

Institutions

Publication Details

Journal
Journal of King Saud University - Computer and Information Sciences
Published
2026-08-26
DOI
https://doi.org/10.1007/s44443-026-01198-0
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Intrinsic modality disparity augmentation without database via cross-modal routing

Hua Wu, Yanxiong Wu, Haotian Hong, Yang Zhou et al.
Journal of King Saud University - Computer and Information Sciences
Multimodal Machine Learning Applications
article

Intrinsic modality disparity augmentation without database via cross-modal routing

Hua Wu, Yanxiong Wu, Haotian Hong, Yang Zhou, Yongfei Zhang, Yuhan Xu, Xiaojing Bai
article en

Abstract

Current augmentation of Large Language Models (LLMs) primarily relies on external corpora, which constrains performance based on database availability and leaves models vulnerable to input perturbations. Inspired by the human perceptual system’s ability to leverage mutual augmentation across different modalities, We propose Intrinsic Modality Disparity Augmentation (IMD), a reliability-aware cross-modal re-encoding framework that does not rely on external retrieval databases during inference, while using a labeled calibration set to estimate modality-specific routing centroids. Our approach consists of two core components: first, an on-demand synthesis mechanism that assesses the reliability of the initial query and generates complementary modal representations when text is deemed insufficient; second, a dynamic cross-modal latent space routing module that evaluates and ranks candidates based on their proximity to learned expert centroids. Extensive experiments on diverse reasoning benchmarks demonstrate that IMD significantly outperforms the MLLM’s native text-only inference and standard augmentation methods, delivering substantial gains in both reasoning accuracy and robustness against textual noise. Furthermore, efficiency analysis reveals that IMD maintains a favorable balance between performance and computational overhead, achieving 3x speedups over naive multimodal conversion and remaining highly competitive with traditional text-based augmentation methods.

Journal of King Saud University - Computer and Information SciencesVol. 38(7)
Hebei University of Engineering (CN), North China Electric Power University (CN), Hebei University of Technology (CN), Hebei Seismological Bureau (CN), Beihang University (CN)
Openalex Percentile: Top 12%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.