Cross-modal audio-text attention for multimodal multitask speech emotion recognition in low-resource Urdu

Speech Emotion Recognition (SER) in low-resource languages remains challenging due to limited annotated data, speaker variability, and the multimodal nature of emotional expression. This paper repositions established components wav2vec 2.0, XLM-R, cross-modal attention, and multitask affective modeling into a framework jointly validated across speaker-independent, cross-lingual zero-shot, and attribution-faithfulness generalization for Urdu, a combination not jointly reported in prior Urdu SER work. The proposed multimodal multitask model achieves 91.3% emotion recognition accuracy on the Urdu Speech Emotion Corpus (UrSEC), outperforming strong audio-only and text-only baselines, with joint valence-arousal learning consistently improving over emotion-only training. Speaker-independent evaluation shows a performance drop relative to random-split testing but confirms substantial robustness to speaker-specific bias. Cross-lingual zero-shot evaluation on English datasets yields 80.2 ± 1.3% (IEMOCAP) to 86.3 ± 0.9% (CREMA-D) accuracy without fine-tuning, indicating substantial cross-lingual transfer, though this alone does not establish full language-neutrality. All performance gains are statistically validated across multiple runs, and attention/Integrated Gradients analyses, supported by quantitative faithfulness testing, show the model relies on emotionally salient acoustic regions and Urdu tokens rather than spurious correlations.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-04
DOI
https://doi.org/10.1038/s41598-026-69201-2
Primary Topic
Emotion and Mood Recognition
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Cross-modal audio-text attention for multimodal multitask speech emotion recognition in low-resource Urdu

Grigori Sidorov, Olga Kolesnikova, Abdullah Abdullah, Zulaikha Fatima et al.
Scientific Reports
Emotion and Mood Recognition
article

Cross-modal audio-text attention for multimodal multitask speech emotion recognition in low-resource Urdu

Grigori Sidorov, Olga Kolesnikova, Abdullah Abdullah, Zulaikha Fatima, Muhammad Ateeb Ather
article en

Abstract

Speech Emotion Recognition (SER) in low-resource languages remains challenging due to limited annotated data, speaker variability, and the multimodal nature of emotional expression. This paper repositions established components wav2vec 2.0, XLM-R, cross-modal attention, and multitask affective modeling into a framework jointly validated across speaker-independent, cross-lingual zero-shot, and attribution-faithfulness generalization for Urdu, a combination not jointly reported in prior Urdu SER work. The proposed multimodal multitask model achieves 91.3% emotion recognition accuracy on the Urdu Speech Emotion Corpus (UrSEC), outperforming strong audio-only and text-only baselines, with joint valence-arousal learning consistently improving over emotion-only training. Speaker-independent evaluation shows a performance drop relative to random-split testing but confirms substantial robustness to speaker-specific bias. Cross-lingual zero-shot evaluation on English datasets yields 80.2 ± 1.3% (IEMOCAP) to 86.3 ± 0.9% (CREMA-D) accuracy without fine-tuning, indicating substantial cross-lingual transfer, though this alone does not establish full language-neutrality. All performance gains are statistically validated across multiple runs, and attention/Integrated Gradients analyses, supported by quantitative faithfulness testing, show the model relies on emotionally salient acoustic regions and Urdu tokens rather than spurious correlations.

Scientific Reports
Superior University (PK), Bahria University (PK), Instituto Politécnico Nacional (MX)
Quality Education
Openalex Percentile: Top 12%
Emotion and Mood Recognition
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Cross-modal audio-text attention for multimodal multitask speech emotion recognition in low-resource Urdu — Grigori Sidorov, Olga Kolesnikova, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS