Multilingual code-switched spoofed speech detection using self-supervised learning models

The rapid advancement of audio generation technologies has enabled the creation of highly realistic synthetic speech, underscoring the need for robust audio spoof detection techniques. However, most existing countermeasures are evaluated on monolingual datasets with utterances from native speakers, leaving their performance on code-switched multilingual speech underexplored. To address this gap, we introduce the multilingual spoofed speech (MSS) dataset, featuring bona fide and spoofed utterances in Urdu, Hindi, and English that capture the natural conversational and code-switching patterns of South Asian speakers. A fused self-supervised learning (SSL) embedding-based audio anti-spoofing framework using an integrated spectro-temporal graph attention network (AASIST) is proposed that combines representations extracted from waveform-to-vector 2.0 and universal speech representation with speaker-aware training frontends. Extensive experiments are conducted on the proposed MSS dataset and the benchmark automatic speaker verification spoofing and countermeasures (ASVspoof) 2019 logical access (LA), ASVspoof-2021 LA, and ASVspoof-2021 deepfake (DF) datasets. Cross-lingual and cross-corpora evaluations demonstrate that the proposed anti-spoofing framework achieves excellent generalization, attaining equal error rates of 0.38%, 0.99%, and 2.46% on ASVspoof-2019 LA, ASVspoof-2021 LA, and ASVspoof-2021 DF, respectively. Moreover, multilingual training on MSS significantly improves the robustness of the implemented artificial intelligence (AI) model, particularly in handling code-switched and linguistically diverse audio. Overall, the results highlight both the engineering contribution of the MSS dataset and the AI contribution of the proposed multi-SSL AASIST framework, together advancing the application of AI for cross-lingual generalizability and the reliable detection of spoofed speech across monolingual and code-switched scenarios.

Authors

Institutions

Publication Details

Journal
Engineering Applications of Artificial Intelligence
Published
2026-10-03
DOI
https://doi.org/10.1016/j.engappai.2026.116398
Primary Topic
Speech Recognition and Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multilingual code-switched spoofed speech detection using self-supervised learning models

Ali Javed, Shah Nawaz, Hafsa Ilyas, Muhammad Haroon Yousaf et al.
Engineering Applications of Artificial Intelligence
Speech Recognition and Synthesis
article

Multilingual code-switched spoofed speech detection using self-supervised learning models

Ali Javed, Shah Nawaz, Hafsa Ilyas, Muhammad Haroon Yousaf, Junaid Mir, Habib Ullah Manzoor, Ahmed Zoha, Muhammad Hamza
article en

Abstract

The rapid advancement of audio generation technologies has enabled the creation of highly realistic synthetic speech, underscoring the need for robust audio spoof detection techniques. However, most existing countermeasures are evaluated on monolingual datasets with utterances from native speakers, leaving their performance on code-switched multilingual speech underexplored. To address this gap, we introduce the multilingual spoofed speech (MSS) dataset, featuring bona fide and spoofed utterances in Urdu, Hindi, and English that capture the natural conversational and code-switching patterns of South Asian speakers. A fused self-supervised learning (SSL) embedding-based audio anti-spoofing framework using an integrated spectro-temporal graph attention network (AASIST) is proposed that combines representations extracted from waveform-to-vector 2.0 and universal speech representation with speaker-aware training frontends. Extensive experiments are conducted on the proposed MSS dataset and the benchmark automatic speaker verification spoofing and countermeasures (ASVspoof) 2019 logical access (LA), ASVspoof-2021 LA, and ASVspoof-2021 deepfake (DF) datasets. Cross-lingual and cross-corpora evaluations demonstrate that the proposed anti-spoofing framework achieves excellent generalization, attaining equal error rates of 0.38%, 0.99%, and 2.46% on ASVspoof-2019 LA, ASVspoof-2021 LA, and ASVspoof-2021 DF, respectively. Moreover, multilingual training on MSS significantly improves the robustness of the implemented artificial intelligence (AI) model, particularly in handling code-switched and linguistically diverse audio. Overall, the results highlight both the engineering contribution of the MSS dataset and the AI contribution of the proposed multi-SSL AASIST framework, together advancing the application of AI for cross-lingual generalizability and the reliable detection of spoofed speech across monolingual and code-switched scenarios.

Engineering Applications of Artificial IntelligenceVol. 184
Johannes Kepler University of Linz (AT), University of Engineering and Technology Taxila (PK), University of Glasgow (GB)
Openalex Percentile: Top 9%
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.