Dynamic adaptive fusion and adversarial alignment model for robust multimodal affective perception in industrial IoT social robots
To achieve natural and robust human–robot interaction, social robots are increasingly serving as key platforms for perception and interaction in smart application scenarios driven by the Industrial Internet of Things (IIoT), such as collaborative smart manufacturing environments, human–robot co-working production lines, and assistive decision-making tasks in complex, open environments. However, within dynamic, open physical environments, a robot's perceptual modalities are frequently partially compromised due to factors such as sensor occlusion, environmental background noise, lighting variations, or non-cooperative user behaviour. A marked decline has been observed in the subject's emotional intelligence, which is indicative of a significant deterioration. The core challenge of deploying robotic affective intelligence is addressed in this paper through the proposal of a dynamic adaptive multimodal affective perception model for social robots. The model combines the Transformer architecture with adversarial feature alignment techniques, enhancing robustness by simulating uncertain omissions in robotic perception. The model's core is designed to capture intermodal dependencies through a multi-granularity cross-modal interaction module. The model employs an adversarial domain alignment module to reduce feature distribution discrepancies between complete and incomplete perception states. Furthermore, a dynamic feature integration mechanism adaptively assesses the reliability of each modality and adjusts fusion weights. The experimental findings, derived from three publicly available datasets (namely, CMU-MOSI, CMU-MOSEI, and IEMOCAP), demonstrate that the model attains sentiment recognition accuracies of 84.2%, 83.7%, and 85.1%, respectively, under conditions of modality deficiency. This performance surpasses that of prevailing mainstream approaches, including early fusion, late fusion, tensor fusion, and multimodal transformers. This provides an effective solution for incomplete multimodal affective analysis in practical applications. Furthermore, this study presents a robust and reliable multimodal affective computing framework for intelligent human–robot interaction within IIoT environments, while also offering an effective approach to affective modeling under conditions of incomplete multimodal perception.
Authors
- Lü Jin (ORCID: https://orcid.org/0000-0001-6408-3659)
- Meifen Chen
- Ji Li
Institutions
- Shenzhen Polytechnic University (CN)
Publication Details
- Journal
- Discover Applied Sciences
- Published
- 2026-09-16
- DOI
- https://doi.org/10.1007/s42452-026-09509-w
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00