Multilingual code-switched spoofed speech detection using self-supervised learning models
The rapid advancement of audio generation technologies has enabled the creation of highly realistic synthetic speech, underscoring the need for robust audio spoof detection techniques. However, most existing countermeasures are evaluated on monolingual datasets with utterances from native speakers, leaving their performance on code-switched multilingual speech underexplored. To address this gap, we introduce the multilingual spoofed speech (MSS) dataset, featuring bona fide and spoofed utterances in Urdu, Hindi, and English that capture the natural conversational and code-switching patterns of South Asian speakers. A fused self-supervised learning (SSL) embedding-based audio anti-spoofing framework using an integrated spectro-temporal graph attention network (AASIST) is proposed that combines representations extracted from waveform-to-vector 2.0 and universal speech representation with speaker-aware training frontends. Extensive experiments are conducted on the proposed MSS dataset and the benchmark automatic speaker verification spoofing and countermeasures (ASVspoof) 2019 logical access (LA), ASVspoof-2021 LA, and ASVspoof-2021 deepfake (DF) datasets. Cross-lingual and cross-corpora evaluations demonstrate that the proposed anti-spoofing framework achieves excellent generalization, attaining equal error rates of 0.38%, 0.99%, and 2.46% on ASVspoof-2019 LA, ASVspoof-2021 LA, and ASVspoof-2021 DF, respectively. Moreover, multilingual training on MSS significantly improves the robustness of the implemented artificial intelligence (AI) model, particularly in handling code-switched and linguistically diverse audio. Overall, the results highlight both the engineering contribution of the MSS dataset and the AI contribution of the proposed multi-SSL AASIST framework, together advancing the application of AI for cross-lingual generalizability and the reliable detection of spoofed speech across monolingual and code-switched scenarios.
Authors
- Ali Javed (ORCID: https://orcid.org/0000-0002-1290-1477)
- Shah Nawaz (ORCID: https://orcid.org/0000-0002-7715-4409)
- Hafsa Ilyas (ORCID: https://orcid.org/0000-0002-2521-7711)
- Muhammad Haroon Yousaf (ORCID: https://orcid.org/0000-0001-8255-1145)
- Junaid Mir (ORCID: https://orcid.org/0000-0002-4587-5121)
- Habib Ullah Manzoor (ORCID: https://orcid.org/0000-0003-0192-7353)
- Ahmed Zoha (ORCID: https://orcid.org/0000-0001-7497-9336)
- Muhammad Hamza
Institutions
- Johannes Kepler University of Linz (AT)
- University of Engineering and Technology Taxila (PK)
- University of Glasgow (GB)
Publication Details
- Journal
- Engineering Applications of Artificial Intelligence
- Published
- 2026-10-03
- DOI
- https://doi.org/10.1016/j.engappai.2026.116398
- Primary Topic
- Speech Recognition and Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00