A Quantum-Inspired Residual Self-Attention Network for Multimodal Sentiment Analysis
Multimodal sentiment analysis is challenging because textual, visual, and acoustic evidence is heterogeneous and oft en weakly aligned. Here, we present QRSAN, a quantum-inspired residual self-attention network that integrates an LSTM text encoder, modality-specific multilayer perceptrons for visual and acoustic inputs, normalized complex-valued encoding, a TFN-derived interaction expansion, real-valued residual self-attention, and classical projection-score mapping. QRSAN runs entirely on classical hardware and does not perform physical quantum computation. Across CMU-MOSI, CMU-MOSEI, and IEMOCAP, QRSAN was evaluated using a common protocol. It achieved the highest numerical mean ACC and Binary_F1 among the evaluated models on CMU-MOSI, whereas its IEMOCAP label-wise accuracy was below that of EF-LSTM. These findings support the utility of combining constrained complex-valued representations with residual self-attention using the evaluated settings, without claiming universal state-of-the-art performance.
Authors
- Yupeng Liu (ORCID: https://orcid.org/0000-0002-8437-6894)
- Xianjie Feng
- Yewang Zhong (ORCID: https://orcid.org/0009-0009-3530-8842)
Institutions
- Harbin University of Science and Technology (CN)
Publication Details
- Journal
- Entropy
- Published
- 2026-09-11
- DOI
- https://doi.org/10.3390/e28091014
- Primary Topic
- Machine Learning in Materials Science
- Type
- article
- Field-Weighted Citation Impact
- 0.00