A dual-domain guided emotion-specialized swin transformer for enhancing speech emotion recognition

Existing Transformer-based speech emotion recognition methods primarily focus on capturing intra-correlations of a sequence through multi-head self-attention mechanisms, or on capturing local key emotional information using windowed attention. However, they neglect the aggregation of inter-relationships between different local regions, which constitutes important cues for speech emotional representation. To address this limitation, this paper proposes a dual-domain guided emotion-specialized Swin Transformer (DGEST), which can jointly identify emotionally salient regions in both temporal and frequency domains and establish emotional dependency interactions between regions. The proposed method first employs a heterogeneous dual-branch architecture to independently encode temporal and frequency domain features. Subsequently, it partitions time–frequency regions according to the inherent spectral distribution variations of emotional expressions. Finally, it achieves dual-domain emotional salient region focusing, and inter-regional dependency relationships through a hierarchical attention mechanism. Extensive experiments on public datasets IEMOCAP, CASIA, and EMODB demonstrate the superiority and efficiency of the proposed DGEST method. This research provides novel insights for dual-domain modeling in robust emotion perception and offers a practical and advanced technical solution for speech emotion recognition tasks.

Authors

Institutions

Publication Details

Journal
Biomedical Signal Processing and Control
Published
2026-09-19
DOI
https://doi.org/10.1016/j.bspc.2026.111503
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A dual-domain guided emotion-specialized swin transformer for enhancing speech emotion recognition

Heming Huang, Yonghong Fan, Feipeng Da
Biomedical Signal Processing and Control
Emotion and Mood Recognition
article

A dual-domain guided emotion-specialized swin transformer for enhancing speech emotion recognition

Heming Huang, Yonghong Fan, Feipeng Da
article en

Abstract

Existing Transformer-based speech emotion recognition methods primarily focus on capturing intra-correlations of a sequence through multi-head self-attention mechanisms, or on capturing local key emotional information using windowed attention. However, they neglect the aggregation of inter-relationships between different local regions, which constitutes important cues for speech emotional representation. To address this limitation, this paper proposes a dual-domain guided emotion-specialized Swin Transformer (DGEST), which can jointly identify emotionally salient regions in both temporal and frequency domains and establish emotional dependency interactions between regions. The proposed method first employs a heterogeneous dual-branch architecture to independently encode temporal and frequency domain features. Subsequently, it partitions time–frequency regions according to the inherent spectral distribution variations of emotional expressions. Finally, it achieves dual-domain emotional salient region focusing, and inter-regional dependency relationships through a hierarchical attention mechanism. Extensive experiments on public datasets IEMOCAP, CASIA, and EMODB demonstrate the superiority and efficiency of the proposed DGEST method. This research provides novel insights for dual-domain modeling in robust emotion perception and offers a practical and advanced technical solution for speech emotion recognition tasks.

Biomedical Signal Processing and ControlVol. 129
Qinghai Normal University (CN), Qinghai Tibetan Hospital (CN), Southeast University (CN)
Openalex Percentile: Top 7%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A dual-domain guided emotion-specialized swin transformer for enhancing speech emotion recognition — Heming Huang, Yonghong Fan, et al. · Biomedical Signal Processing and Control (2026) | TGRS Research Map | TGRS