Dynamic hypernode injection and graph attention fusion for multimodal sentiment analysis of behavioral signals

Language use, vocal behavior, and facial activity provide complementary indicators of affective state, but multimodal models can be affected by noisy temporal observations and insufficient global guidance before cross-modal interaction. We developed a framework combining bidirectional gated recurrent unit encoders, attention pooling, dynamic hypernode injection, and graph attention fusion. Textual, acoustic, and visual sequences were mapped into a shared latent space and compressed into modality-level representations. A sample-dependent hypernode and a learnable static prior were then injected through gated residual connections before graph propagation. The model was evaluated on CMU-MOSI and CMU-MOSEI using five random seeds and validation-MAE checkpoint selection. On CMU-MOSI, the model obtained an MAE of 0.7290 ± 0.0099 , a correlation of 0.7880 ± 0.0076 , and an Acc-7 of 0.4561 ± 0.0160 . On CMU-MOSEI, the corresponding results were 0.5541 ± 0.0053 , 0.7551 ± 0.0058 , and 0.5269 ± 0.0064 . A five-run comparison of pre-GAT, post-GAT, and no injection showed that pre-GAT injection provided the strongest overall regression-oriented balance, although post-GAT was slightly higher on MOSEI Acc-7. Gate, graph-attention, representation, and balanced case diagnostics showed that the injected prior was sample dependent and modality selective. These findings support pre-propagation hypernode injection under complete, word-aligned tri-modal inputs, without establishing universal superiority or robustness to missing and corrupted modalities.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-28
DOI
https://doi.org/10.1371/journal.pone.0359425
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Dynamic hypernode injection and graph attention fusion for multimodal sentiment analysis of behavioral signals

Yinghao Li, Huizi Yan, Zhuowei Hu, Jian Xu et al.
PLoS ONE
Emotion and Mood Recognition
article

Dynamic hypernode injection and graph attention fusion for multimodal sentiment analysis of behavioral signals

Yinghao Li, Huizi Yan, Zhuowei Hu, Jian Xu, Wenxu Chen, Wenfei Cao
article en

Abstract

Language use, vocal behavior, and facial activity provide complementary indicators of affective state, but multimodal models can be affected by noisy temporal observations and insufficient global guidance before cross-modal interaction. We developed a framework combining bidirectional gated recurrent unit encoders, attention pooling, dynamic hypernode injection, and graph attention fusion. Textual, acoustic, and visual sequences were mapped into a shared latent space and compressed into modality-level representations. A sample-dependent hypernode and a learnable static prior were then injected through gated residual connections before graph propagation. The model was evaluated on CMU-MOSI and CMU-MOSEI using five random seeds and validation-MAE checkpoint selection. On CMU-MOSI, the model obtained an MAE of 0.7290 ± 0.0099 , a correlation of 0.7880 ± 0.0076 , and an Acc-7 of 0.4561 ± 0.0160 . On CMU-MOSEI, the corresponding results were 0.5541 ± 0.0053 , 0.7551 ± 0.0058 , and 0.5269 ± 0.0064 . A five-run comparison of pre-GAT, post-GAT, and no injection showed that pre-GAT injection provided the strongest overall regression-oriented balance, although post-GAT was slightly higher on MOSEI Acc-7. Gate, graph-attention, representation, and balanced case diagnostics showed that the injected prior was sample dependent and modality selective. These findings support pre-propagation hypernode injection under complete, word-aligned tri-modal inputs, without establishing universal superiority or robustness to missing and corrupted modalities.

PLoS ONEVol. 21(9)
Hubei University for Nationalities (CN)
Quality Education
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Dynamic hypernode injection and graph attention fusion for multimodal sentiment analysis of behavioral signals — Yinghao Li, Huizi Yan, et al. · PLoS ONE (2026) | TGRS Research Map | TGRS