Reliability-Conditioned Heterogeneous Foundation-Model Residual Fusion for Multimodal Disaster Classification

During sudden-onset disasters, social media text and images provide time-critical evidence for emergency awareness, damage assessment, and humanitarian response, but their cues are often incomplete, noisy, or conflicting. Direct equal-status fusion of heterogeneous pretrained representations can introduce semantic misalignment and allow unreliable evidence to influence the classifier. This paper proposes foundation-augmented reliability-conditioned dynamic adaptive fusion (FA–RC–DAF) for multimodal disaster classification. The model first preserves CLIP as a unit-weight text–image alignment base, maintaining a stable cross-modal decision space. It then projects BERTweet and SigLIP features into the CLIP-aligned space as bounded residual corrections, enabling domain-specific linguistic and complementary vision–language cues to refine the base without overwriting it. A confidence–agreement router estimates sample-level residual reliability from normalized-entropy predictive concentration and cross-encoder agreement, selectively admitting each residual before fusion. Explicit cross-modal interaction is performed only after the two streams have been reliability-refined. On CrisisMMD, FA–RC–DAF achieves 92.26% accuracy, 92.25% weighted F1, and 91.26% macro F1 under the retained five-class protocol. The protocol-aware published comparison is reported separately from the five-seed internally matched architectural comparison, which provides the primary evidence for method-level claims. Additional evaluations show differentiated behavior under event- and disaster-type shifts, stronger difficulty under forward temporal drift, and condition-dependent sensitivity to corrupted or unavailable inputs, providing a more fully characterized basis for multimodal disaster decision support.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-16
DOI
https://doi.org/10.3390/electronics15184222
Primary Topic
Public Relations and Crisis Communication
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reliability-Conditioned Heterogeneous Foundation-Model Residual Fusion for Multimodal Disaster Classification

Qingjie Liu, Zhian Pan, Guan Li, Lingfeng Niu et al.
Electronics
Public Relations and Crisis Communication
article

Reliability-Conditioned Heterogeneous Foundation-Model Residual Fusion for Multimodal Disaster Classification

Qingjie Liu, Zhian Pan, Guan Li, Lingfeng Niu, Shanshan Li
article en

Abstract

During sudden-onset disasters, social media text and images provide time-critical evidence for emergency awareness, damage assessment, and humanitarian response, but their cues are often incomplete, noisy, or conflicting. Direct equal-status fusion of heterogeneous pretrained representations can introduce semantic misalignment and allow unreliable evidence to influence the classifier. This paper proposes foundation-augmented reliability-conditioned dynamic adaptive fusion (FA–RC–DAF) for multimodal disaster classification. The model first preserves CLIP as a unit-weight text–image alignment base, maintaining a stable cross-modal decision space. It then projects BERTweet and SigLIP features into the CLIP-aligned space as bounded residual corrections, enabling domain-specific linguistic and complementary vision–language cues to refine the base without overwriting it. A confidence–agreement router estimates sample-level residual reliability from normalized-entropy predictive concentration and cross-encoder agreement, selectively admitting each residual before fusion. Explicit cross-modal interaction is performed only after the two streams have been reliability-refined. On CrisisMMD, FA–RC–DAF achieves 92.26% accuracy, 92.25% weighted F1, and 91.26% macro F1 under the retained five-class protocol. The protocol-aware published comparison is reported separately from the five-seed internally matched architectural comparison, which provides the primary evidence for method-level claims. Additional evaluations show differentiated behavior under event- and disaster-type shifts, stronger difficulty under forward temporal drift, and condition-dependent sensitivity to corrupted or unavailable inputs, providing a more fully characterized basis for multimodal disaster decision support.

ElectronicsVol. 15(18)
China People's Public Security University (CN), Beijing Institute of Big Data Research (CN), China Information Technology Security Evaluation Center (CN), Beijing Information Science & Technology University (CN)
Climate action
Openalex Percentile: Top 4%
Public Relations and Crisis Communication
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.