Towards inclusive voice biometrics: Dysarthria-discriminative embeddings for ASV system
Automatic speaker verification (ASV) performance degrades considerably for individuals with dysarthria due to articulatory impairments and increased acoustic variability. This study proposes two discriminative front-end acoustic feature embeddings to improve dysarthric ASV: discriminative bottleneck (DBN) embeddings learned through a CNN-based control-dysarthria classification task, and discriminative contrastive bottleneck embeddings (DCBN) derived from a Siamese network with contrastive learning. Each embedding type is independently fused with Mel-frequency cepstral coefficients and refined via principal component analysis to form compact ASV front end representation. Experimental evaluations on the TORGO and UA-Speech datasets demonstrate consistent improvements across all dysarthria severity levels, with DCBN achieving the most substantial gains due to enhanced speaker-discriminative learning. When integrated with an ECAPA-TDNN back end, the proposed approach yields relative reductions of 16.35% in equal error rate (EER) and 31.75% in minimum detection cost function (min DCF) over the baseline. These findings confirm the effectiveness of dysarthria discriminative embeddings in enabling more reliable and inclusive ASV system for dysarthric speakers.
Authors
- S. Shahnawazuddin (ORCID: https://orcid.org/0000-0002-3916-9693)
- Waquar Ahmad (ORCID: https://orcid.org/0000-0001-7817-3313)
- Shinimol Salim
Institutions
- National Institute of Technology Calicut (IN)
- National Institute of Technology Patna (IN)
- Rajamangala University of Technology Isan (TH)
Publication Details
- Journal
- Computers & Electrical Engineering
- Published
- 2026-09-19
- DOI
- https://doi.org/10.1016/j.compeleceng.2026.111531
- Primary Topic
- Voice and Speech Disorders
- Type
- article
- Field-Weighted Citation Impact
- 0.00