Dynamic adaptive fusion and adversarial alignment model for robust multimodal affective perception in industrial IoT social robots

To achieve natural and robust human–robot interaction, social robots are increasingly serving as key platforms for perception and interaction in smart application scenarios driven by the Industrial Internet of Things (IIoT), such as collaborative smart manufacturing environments, human–robot co-working production lines, and assistive decision-making tasks in complex, open environments. However, within dynamic, open physical environments, a robot's perceptual modalities are frequently partially compromised due to factors such as sensor occlusion, environmental background noise, lighting variations, or non-cooperative user behaviour. A marked decline has been observed in the subject's emotional intelligence, which is indicative of a significant deterioration. The core challenge of deploying robotic affective intelligence is addressed in this paper through the proposal of a dynamic adaptive multimodal affective perception model for social robots. The model combines the Transformer architecture with adversarial feature alignment techniques, enhancing robustness by simulating uncertain omissions in robotic perception. The model's core is designed to capture intermodal dependencies through a multi-granularity cross-modal interaction module. The model employs an adversarial domain alignment module to reduce feature distribution discrepancies between complete and incomplete perception states. Furthermore, a dynamic feature integration mechanism adaptively assesses the reliability of each modality and adjusts fusion weights. The experimental findings, derived from three publicly available datasets (namely, CMU-MOSI, CMU-MOSEI, and IEMOCAP), demonstrate that the model attains sentiment recognition accuracies of 84.2%, 83.7%, and 85.1%, respectively, under conditions of modality deficiency. This performance surpasses that of prevailing mainstream approaches, including early fusion, late fusion, tensor fusion, and multimodal transformers. This provides an effective solution for incomplete multimodal affective analysis in practical applications. Furthermore, this study presents a robust and reliable multimodal affective computing framework for intelligent human–robot interaction within IIoT environments, while also offering an effective approach to affective modeling under conditions of incomplete multimodal perception.

Authors

Institutions

Publication Details

Journal
Discover Applied Sciences
Published
2026-09-16
DOI
https://doi.org/10.1007/s42452-026-09509-w
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Dynamic adaptive fusion and adversarial alignment model for robust multimodal affective perception in industrial IoT social robots

Lü Jin, Meifen Chen, Ji Li
Discover Applied Sciences
Emotion and Mood Recognition
article

Dynamic adaptive fusion and adversarial alignment model for robust multimodal affective perception in industrial IoT social robots

Lü Jin, Meifen Chen, Ji Li
article en

Abstract

To achieve natural and robust human–robot interaction, social robots are increasingly serving as key platforms for perception and interaction in smart application scenarios driven by the Industrial Internet of Things (IIoT), such as collaborative smart manufacturing environments, human–robot co-working production lines, and assistive decision-making tasks in complex, open environments. However, within dynamic, open physical environments, a robot's perceptual modalities are frequently partially compromised due to factors such as sensor occlusion, environmental background noise, lighting variations, or non-cooperative user behaviour. A marked decline has been observed in the subject's emotional intelligence, which is indicative of a significant deterioration. The core challenge of deploying robotic affective intelligence is addressed in this paper through the proposal of a dynamic adaptive multimodal affective perception model for social robots. The model combines the Transformer architecture with adversarial feature alignment techniques, enhancing robustness by simulating uncertain omissions in robotic perception. The model's core is designed to capture intermodal dependencies through a multi-granularity cross-modal interaction module. The model employs an adversarial domain alignment module to reduce feature distribution discrepancies between complete and incomplete perception states. Furthermore, a dynamic feature integration mechanism adaptively assesses the reliability of each modality and adjusts fusion weights. The experimental findings, derived from three publicly available datasets (namely, CMU-MOSI, CMU-MOSEI, and IEMOCAP), demonstrate that the model attains sentiment recognition accuracies of 84.2%, 83.7%, and 85.1%, respectively, under conditions of modality deficiency. This performance surpasses that of prevailing mainstream approaches, including early fusion, late fusion, tensor fusion, and multimodal transformers. This provides an effective solution for incomplete multimodal affective analysis in practical applications. Furthermore, this study presents a robust and reliable multimodal affective computing framework for intelligent human–robot interaction within IIoT environments, while also offering an effective approach to affective modeling under conditions of incomplete multimodal perception.

Discover Applied Sciences
Shenzhen Polytechnic University (CN)
Peace, Justice and strong institutions
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.