Evaluating early, late, hybrid and meta fusion in multimodal emotion detection with pretrained models

Understanding emotions in dialogue is essential for socially intelligent agents; however, the challenge in multimodal approaches comes from the combination of multiple modalities of speech, language, and visual expressions. In this paper, we propose an experimental benchmark of four well-known multimodal fusion techniques: early fusion, late fusion, hybrid averaging, and meta fusion for the MELD dataset based on frozen pretrained encoders for multimodal emotion recognition. We use existing pre-trained encoders for text, speech, and a single representative image frame from each utterance to construct complex representations of utterances, and then apply them to four fusion techniques: early fusion, late fusion, hybrid averaging, and a meta classifier. As the visual stream is extracted from one single static frame, there is no temporal dynamics in the facial expressions captured by the visual stream. The results indicate that the proposed evaluation framework achieves better performance compared to strong single-modality baselines, as early, hybrid, and meta fusion yield statistically similar performances, all far outperforming the purely textual baseline, especially when considering strong emotions like anger, joy, and surprise. However, early and meta fusions do not work at all for the fear emotion (F1 = 0.000), while the recognition of disgust is poor overall, with best F1 = 0.220 for hybrid fusion, due to their very low occurrences in the data (1.9% and 2.6%, respectively). Based on the analysis of the performances of each class and the confusion patterns, it is evident that both the hybrid and meta-fusion approaches leverage the best of both the feature-based and the score-based fusion approach while keeping the number of task-specific parameters minimal. The above results form a reproducible benchmark for the comparison of the well-known multimodal fusion approaches and guide one to choose the fusion method in the case of pre-trained encoders. This benchmark provides a practical reference point for selecting fusion strategies in resource-constrained multimodal emotion recognition systems, with applications spanning customer service, healthcare, and education.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-15
DOI
https://doi.org/10.1038/s41598-026-68674-5
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating early, late, hybrid and meta fusion in multimodal emotion detection with pretrained models

Awani Bhushan, Syed Riyas Ahamed, Sandip Saha
Scientific Reports
Emotion and Mood Recognition
article

Evaluating early, late, hybrid and meta fusion in multimodal emotion detection with pretrained models

Awani Bhushan, Syed Riyas Ahamed, Sandip Saha
article en

Abstract

Understanding emotions in dialogue is essential for socially intelligent agents; however, the challenge in multimodal approaches comes from the combination of multiple modalities of speech, language, and visual expressions. In this paper, we propose an experimental benchmark of four well-known multimodal fusion techniques: early fusion, late fusion, hybrid averaging, and meta fusion for the MELD dataset based on frozen pretrained encoders for multimodal emotion recognition. We use existing pre-trained encoders for text, speech, and a single representative image frame from each utterance to construct complex representations of utterances, and then apply them to four fusion techniques: early fusion, late fusion, hybrid averaging, and a meta classifier. As the visual stream is extracted from one single static frame, there is no temporal dynamics in the facial expressions captured by the visual stream. The results indicate that the proposed evaluation framework achieves better performance compared to strong single-modality baselines, as early, hybrid, and meta fusion yield statistically similar performances, all far outperforming the purely textual baseline, especially when considering strong emotions like anger, joy, and surprise. However, early and meta fusions do not work at all for the fear emotion (F1 = 0.000), while the recognition of disgust is poor overall, with best F1 = 0.220 for hybrid fusion, due to their very low occurrences in the data (1.9% and 2.6%, respectively). Based on the analysis of the performances of each class and the confusion patterns, it is evident that both the hybrid and meta-fusion approaches leverage the best of both the feature-based and the score-based fusion approach while keeping the number of task-specific parameters minimal. The above results form a reproducible benchmark for the comparison of the well-known multimodal fusion approaches and guide one to choose the fusion method in the case of pre-trained encoders. This benchmark provides a practical reference point for selecting fusion strategies in resource-constrained multimodal emotion recognition systems, with applications spanning customer service, healthcare, and education.

Scientific Reports
Vellore Institute of Technology University (IN)
Quality Education
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.