A hierarchical multistage semantic fusion framework for multimodal sentiment analysis

Purpose Multimodal sentiment analysis aims to improve cross-modal fusion to understand sentiment better. Most existing methods rely on single-stage fusion or local cross-modal interactions, making it difficult to fully capture relationships among modalities, thereby limiting their sentiment representation capabilities and predictive performance. Design/methodology/approach This study proposes a Hierarchical Multistage Semantic Fusion (HMSF) framework. First, modality-specific encoding and unified projection are employed to achieve cross-modal alignment among textual, acoustic, and visual modalities. Next, a hierarchical multistage fusion structure is introduced to integrate multimodal information progressively. Finally, a gated contextual updating mechanism dynamically aggregates cross-sample contextual information to optimize fused representations. The model is trained with a unified objective for sentiment classification and continuous regression tasks. Findings Experimental results demonstrate that HMSF achieves competitive overall performance compared with representative existing methods on both the CMU-MOSEI and CH-SIMS datasets. On CMU-MOSEI, HMSF improves Acc-7 by approximately 4.60 percentage points over the average performance of the baseline methods. On CH-SIMS, HMSF improves the F1 score by approximately 1.44 percentage points over the average performance of the baseline methods. Originality/value By combining hierarchical fusion with gated contextual updating, HMSF enhances multimodal feature interaction and shows potential for applications such as online content analysis, public opinion monitoring, and user sentiment understanding.

Authors

Institutions

Publication Details

Journal
Aslib Journal of Information Management
Published
2026-09-25
DOI
https://doi.org/10.1108/ajim-01-2026-0042
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A hierarchical multistage semantic fusion framework for multimodal sentiment analysis

Guangyu Mu, Jiaxue Li, Hong-Bo Sun, Jiaxiu Dai
Aslib Journal of Information Management
Emotion and Mood Recognition
article

A hierarchical multistage semantic fusion framework for multimodal sentiment analysis

Guangyu Mu, Jiaxue Li, Hong-Bo Sun, Jiaxiu Dai
article en

Abstract

Purpose Multimodal sentiment analysis aims to improve cross-modal fusion to understand sentiment better. Most existing methods rely on single-stage fusion or local cross-modal interactions, making it difficult to fully capture relationships among modalities, thereby limiting their sentiment representation capabilities and predictive performance. Design/methodology/approach This study proposes a Hierarchical Multistage Semantic Fusion (HMSF) framework. First, modality-specific encoding and unified projection are employed to achieve cross-modal alignment among textual, acoustic, and visual modalities. Next, a hierarchical multistage fusion structure is introduced to integrate multimodal information progressively. Finally, a gated contextual updating mechanism dynamically aggregates cross-sample contextual information to optimize fused representations. The model is trained with a unified objective for sentiment classification and continuous regression tasks. Findings Experimental results demonstrate that HMSF achieves competitive overall performance compared with representative existing methods on both the CMU-MOSEI and CH-SIMS datasets. On CMU-MOSEI, HMSF improves Acc-7 by approximately 4.60 percentage points over the average performance of the baseline methods. On CH-SIMS, HMSF improves the F1 score by approximately 1.44 percentage points over the average performance of the baseline methods. Originality/value By combining hierarchical fusion with gated contextual updating, HMSF enhances multimodal feature interaction and shows potential for applications such as online content analysis, public opinion monitoring, and user sentiment understanding.

Aslib Journal of Information Management
Jilin University of Finance and Economics (CN), Jilin University (CN), Jilin Medical University (CN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A hierarchical multistage semantic fusion framework for multimodal sentiment analysis — Guangyu Mu, Jiaxue Li, et al. · Aslib Journal of Information Management (2026) | TGRS Research Map | TGRS