QESM: A Quantum-Inspired Network with Dynamic Fusion and Entangled Measurement for multimodal emotion recognition

Multimodal emotion recognition in conversation (MERC) integrates textual, acoustic, visual, and contextual cues to identify each utterance’s emotion. Classical models encode temporal and cross-modal relations through recurrence, graphs, or attention, whereas quantum-inspired models largely retain static states and separable measurements, leaving operator-level evolution and coherence-aware joint readout under-modelled. To address these limitations, we propose QESM, which integrates Unitary Temporal Evolution (UTE), OTOC Cross-Modal Scrambling (OCS), and Entangled Born Measurement (EBM). UTE propagates speaker-aware emotional context through norm-preserving unitary phase evolution. Inspired by the OTOC, OCS models dynamic cross-modal complementarity by quantifying how strongly one temporally evolved modality perturbs the interpretation of another. Inspired by the non-separability of quantum entanglement, EBM constructs a bipartite entangled state over the modality-identity and modality-representation spaces and separately derives class-wise Born measurement responses from cross-modal phase coherence and trimodal class agreement. Experiments on the 7433-utterance IEMOCAP and 13,708-utterance MELD datasets show that QESM achieves the highest weighted-F1 among the compared methods, reaching 73.23% and 67.31%, respectively, with five-seed means of 72.82 ± 0.23 % and 67.17 ± 0.11 % . Based on the unrounded WF1 scores, these representative results are 1.29 and 1.08 percentage points higher than QR-Net, respectively. Paired tests show significant gains over several baselines, whereas the difference between QESM and QR-Net is nonsignificant after Holm correction. Ablation and control studies show that QESM effectively captures temporal context and dynamic cross-modal complementarity while enabling coherence-aware joint measurement

Authors

Institutions

Publication Details

Journal
Information Processing & Management
Published
2026-09-12
DOI
https://doi.org/10.1016/j.ipm.2026.105163
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

QESM: A Quantum-Inspired Network with Dynamic Fusion and Entangled Measurement for multimodal emotion recognition

Hong Zhao, Yuqi Wang, Yibin Guo, Jiahao Wu et al.
Information Processing & Management
Emotion and Mood Recognition
article

QESM: A Quantum-Inspired Network with Dynamic Fusion and Entangled Measurement for multimodal emotion recognition

Hong Zhao, Yuqi Wang, Yibin Guo, Jiahao Wu, Zhicheng Huang, Yangcan Liao
article en

Abstract

Multimodal emotion recognition in conversation (MERC) integrates textual, acoustic, visual, and contextual cues to identify each utterance’s emotion. Classical models encode temporal and cross-modal relations through recurrence, graphs, or attention, whereas quantum-inspired models largely retain static states and separable measurements, leaving operator-level evolution and coherence-aware joint readout under-modelled. To address these limitations, we propose QESM, which integrates Unitary Temporal Evolution (UTE), OTOC Cross-Modal Scrambling (OCS), and Entangled Born Measurement (EBM). UTE propagates speaker-aware emotional context through norm-preserving unitary phase evolution. Inspired by the OTOC, OCS models dynamic cross-modal complementarity by quantifying how strongly one temporally evolved modality perturbs the interpretation of another. Inspired by the non-separability of quantum entanglement, EBM constructs a bipartite entangled state over the modality-identity and modality-representation spaces and separately derives class-wise Born measurement responses from cross-modal phase coherence and trimodal class agreement. Experiments on the 7433-utterance IEMOCAP and 13,708-utterance MELD datasets show that QESM achieves the highest weighted-F1 among the compared methods, reaching 73.23% and 67.31%, respectively, with five-seed means of 72.82 ± 0.23 % and 67.17 ± 0.11 % . Based on the unrounded WF1 scores, these representative results are 1.29 and 1.08 percentage points higher than QR-Net, respectively. Paired tests show significant gains over several baselines, whereas the difference between QESM and QR-Net is nonsignificant after Holm correction. Ablation and control studies show that QESM effectively captures temporal context and dynamic cross-modal complementarity while enabling coherence-aware joint measurement

Information Processing & ManagementVol. 64(2)
Fuzhou University (CN), Minnan Normal University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.