Empirical mode decomposition based CNN prediction model for speech emotion classification using Mel scale spectrograms
Abstract Studies on emotion recognition systems from speech signal are comprehensively assessed in the purview of developing Human-Computer Interface (HCI) models. Traditionally, HCI models offer significant contribution into the psychological states of individuals, serving as crucial psychic response for HCI applications. Majority of the existing works in literature concentrate on classifying emotional states with several spectral and Cepstrum features using ANN and RNN based classification for various data sets. Despite that, the results are found to be comparatively low for RAVDESS dataset. Furthermore, only limited studies focus on using 3D time frequency images for predicting Speech Emotion Recognition (SER) models. In light of these aspects, the proposed work aims at recognizing the speech emotion with seven different states using deep neural net with Empirical mode decomposition (EMD).In the initial phase, the raw speech signals are decomposed into 5 different Intrinsic Mode Functions (IMFs) and are analysed in time, frequency and energy. Subsequently, all the five oscillatory modes are transformed into linear scale and Mel scale spectrograms using short time frequency transform and auditory filter banks. These spectrograms are fed as input to three different pre-trained models namely AlexNet, GoogLeNet and ResNet-50 and evaluated for all IMFs both in linear and Mel Scale. Ablation studies are also carried out with controlled baselines viz. Mel/ Linear scale spectrograms without EMD and classical Machine Learning classifiers for fair comparison. The experimental results record an improved performance in accuracy for IMF1 + Mel scale spectrograms + ResNet-50 classifier with 76.38% and linear scale of 74.25%. Additionally, 5- fold cross validation analysis is also carried out for generalizing the model and the results are reported. In comparison to the other approaches used in literature for categorizing RAVDESS dataset, the proposed method found to yield better results.
Authors
- S. Jayalakshmy (ORCID: https://orcid.org/0000-0001-8876-239X)
- Lakshmipriya Balagouruchetty (ORCID: https://orcid.org/0000-0002-4460-6915)
Institutions
- Jawaharlal Nehru University (IN)
Publication Details
- Journal
- Journal of Intelligent & Fuzzy Systems
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1177/18758967261476147
- Primary Topic
- Emotion and Mood Recognition
- Type
- article
- Field-Weighted Citation Impact
- 0.00