Evaluating early, late, hybrid and meta fusion in multimodal emotion detection with pretrained models
Understanding emotions in dialogue is essential for socially intelligent agents; however, the challenge in multimodal approaches comes from the combination of multiple modalities of speech, language, and visual expressions. In this paper, we propose an experimental benchmark of four well-known multimodal fusion techniques: early fusion, late fusion, hybrid averaging, and meta fusion for the MELD dataset based on frozen pretrained encoders for multimodal emotion recognition. We use existing pre-trained encoders for text, speech, and a single representative image frame from each utterance to construct complex representations of utterances, and then apply them to four fusion techniques: early fusion, late fusion, hybrid averaging, and a meta classifier. As the visual stream is extracted from one single static frame, there is no temporal dynamics in the facial expressions captured by the visual stream. The results indicate that the proposed evaluation framework achieves better performance compared to strong single-modality baselines, as early, hybrid, and meta fusion yield statistically similar performances, all far outperforming the purely textual baseline, especially when considering strong emotions like anger, joy, and surprise. However, early and meta fusions do not work at all for the fear emotion (F1 = 0.000), while the recognition of disgust is poor overall, with best F1 = 0.220 for hybrid fusion, due to their very low occurrences in the data (1.9% and 2.6%, respectively). Based on the analysis of the performances of each class and the confusion patterns, it is evident that both the hybrid and meta-fusion approaches leverage the best of both the feature-based and the score-based fusion approach while keeping the number of task-specific parameters minimal. The above results form a reproducible benchmark for the comparison of the well-known multimodal fusion approaches and guide one to choose the fusion method in the case of pre-trained encoders. This benchmark provides a practical reference point for selecting fusion strategies in resource-constrained multimodal emotion recognition systems, with applications spanning customer service, healthcare, and education.
Authors
- Awani Bhushan (ORCID: https://orcid.org/0000-0002-4150-5406)
- Syed Riyas Ahamed
- Sandip Saha
Institutions
- Vellore Institute of Technology University (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1038/s41598-026-68674-5
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00