Empirical mode decomposition based CNN prediction model for speech emotion classification using Mel scale spectrograms

Abstract Studies on emotion recognition systems from speech signal are comprehensively assessed in the purview of developing Human-Computer Interface (HCI) models. Traditionally, HCI models offer significant contribution into the psychological states of individuals, serving as crucial psychic response for HCI applications. Majority of the existing works in literature concentrate on classifying emotional states with several spectral and Cepstrum features using ANN and RNN based classification for various data sets. Despite that, the results are found to be comparatively low for RAVDESS dataset. Furthermore, only limited studies focus on using 3D time frequency images for predicting Speech Emotion Recognition (SER) models. In light of these aspects, the proposed work aims at recognizing the speech emotion with seven different states using deep neural net with Empirical mode decomposition (EMD).In the initial phase, the raw speech signals are decomposed into 5 different Intrinsic Mode Functions (IMFs) and are analysed in time, frequency and energy. Subsequently, all the five oscillatory modes are transformed into linear scale and Mel scale spectrograms using short time frequency transform and auditory filter banks. These spectrograms are fed as input to three different pre-trained models namely AlexNet, GoogLeNet and ResNet-50 and evaluated for all IMFs both in linear and Mel Scale. Ablation studies are also carried out with controlled baselines viz. Mel/ Linear scale spectrograms without EMD and classical Machine Learning classifiers for fair comparison. The experimental results record an improved performance in accuracy for IMF1 + Mel scale spectrograms + ResNet-50 classifier with 76.38% and linear scale of 74.25%. Additionally, 5- fold cross validation analysis is also carried out for generalizing the model and the results are reported. In comparison to the other approaches used in literature for categorizing RAVDESS dataset, the proposed method found to yield better results.

Authors

Institutions

Publication Details

Journal
Journal of Intelligent & Fuzzy Systems
Published
2026-09-15
DOI
https://doi.org/10.1177/18758967261476147
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Empirical mode decomposition based CNN prediction model for speech emotion classification using Mel scale spectrograms

S. Jayalakshmy, Lakshmipriya Balagouruchetty
Journal of Intelligent & Fuzzy Systems
Emotion and Mood Recognition
article

Empirical mode decomposition based CNN prediction model for speech emotion classification using Mel scale spectrograms

S. Jayalakshmy, Lakshmipriya Balagouruchetty
article en

Abstract

Abstract Studies on emotion recognition systems from speech signal are comprehensively assessed in the purview of developing Human-Computer Interface (HCI) models. Traditionally, HCI models offer significant contribution into the psychological states of individuals, serving as crucial psychic response for HCI applications. Majority of the existing works in literature concentrate on classifying emotional states with several spectral and Cepstrum features using ANN and RNN based classification for various data sets. Despite that, the results are found to be comparatively low for RAVDESS dataset. Furthermore, only limited studies focus on using 3D time frequency images for predicting Speech Emotion Recognition (SER) models. In light of these aspects, the proposed work aims at recognizing the speech emotion with seven different states using deep neural net with Empirical mode decomposition (EMD).In the initial phase, the raw speech signals are decomposed into 5 different Intrinsic Mode Functions (IMFs) and are analysed in time, frequency and energy. Subsequently, all the five oscillatory modes are transformed into linear scale and Mel scale spectrograms using short time frequency transform and auditory filter banks. These spectrograms are fed as input to three different pre-trained models namely AlexNet, GoogLeNet and ResNet-50 and evaluated for all IMFs both in linear and Mel Scale. Ablation studies are also carried out with controlled baselines viz. Mel/ Linear scale spectrograms without EMD and classical Machine Learning classifiers for fair comparison. The experimental results record an improved performance in accuracy for IMF1 + Mel scale spectrograms + ResNet-50 classifier with 76.38% and linear scale of 74.25%. Additionally, 5- fold cross validation analysis is also carried out for generalizing the model and the results are reported. In comparison to the other approaches used in literature for categorizing RAVDESS dataset, the proposed method found to yield better results.

Journal of Intelligent & Fuzzy Systems
Jawaharlal Nehru University (IN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.