DCapsNet: a robust multi-feature deep capsule network for speech emotion recognition

With the continuous advancements in artificial intelligence, speech emotion recognition (SER) systems have garnered widespread attention. They have found applications across various fields, including mobile services, psychological assessment, healthcare, and autonomous vehicles. This research introduces a novel multi-feature deep capsule network capable of real-time speech processing. Firstly, speech preprocessing and data enhancement techniques were used to enhance the dataset, followed by the extraction of discriminative features. To construct the SER model, spectral and temporal information was extracted as a set using three datasets: SAVEE, RAVDESS, and EMO-DB. These feature sets were given to the deep capsule network (DCapsNet). The capsule layers in DCapsNet preserved the spatial relationships within speech signals and ensured robust feature representation, allowing for accurate emotion recognition. The model employed a hierarchical learning using aggregated capsule outputs to enhance recognition performance. Finally, the performance of DCapsNet is assessed and compared with that of the baseline models using the multilingual speech datasets. The results demonstrated that the DCapsNet achieved superior performance compared to existing methods in the field. The micro accuracy of DCapsNet for the RAVDESS, SAVEE, and EMO-DB datasets is 94%, 90%, and 93%, respectively. This study demonstrated the effectiveness of the DCapsNet for recognizing speech emotions.

Authors

Institutions

Publication Details

Journal
Automatika
Published
2026-09-16
DOI
https://doi.org/10.1080/00051144.2026.2727198
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

DCapsNet: a robust multi-feature deep capsule network for speech emotion recognition

Rashid Jahangir, Faiqa Hanif, Nazik Alturki, Mohammed Alreshoodi
Automatika
Emotion and Mood Recognition
article

DCapsNet: a robust multi-feature deep capsule network for speech emotion recognition

Rashid Jahangir, Faiqa Hanif, Nazik Alturki, Mohammed Alreshoodi
article en

Abstract

With the continuous advancements in artificial intelligence, speech emotion recognition (SER) systems have garnered widespread attention. They have found applications across various fields, including mobile services, psychological assessment, healthcare, and autonomous vehicles. This research introduces a novel multi-feature deep capsule network capable of real-time speech processing. Firstly, speech preprocessing and data enhancement techniques were used to enhance the dataset, followed by the extraction of discriminative features. To construct the SER model, spectral and temporal information was extracted as a set using three datasets: SAVEE, RAVDESS, and EMO-DB. These feature sets were given to the deep capsule network (DCapsNet). The capsule layers in DCapsNet preserved the spatial relationships within speech signals and ensured robust feature representation, allowing for accurate emotion recognition. The model employed a hierarchical learning using aggregated capsule outputs to enhance recognition performance. Finally, the performance of DCapsNet is assessed and compared with that of the baseline models using the multilingual speech datasets. The results demonstrated that the DCapsNet achieved superior performance compared to existing methods in the field. The micro accuracy of DCapsNet for the RAVDESS, SAVEE, and EMO-DB datasets is 94%, 90%, and 93%, respectively. This study demonstrated the effectiveness of the DCapsNet for recognizing speech emotions.

AutomatikaVol. 67(1)
Princess Nourah bint Abdulrahman University (SA), Qassim University (SA), COMSATS University Islamabad (PK)
Princess Nourah Bint Abdulrahman University
Gender equality
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

DCapsNet: a robust multi-feature deep capsule network for speech emotion recognition — Rashid Jahangir, Faiqa Hanif, et al. · Automatika (2026) | TGRS Research Map | TGRS