Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network
Abstract Speech-language pathologists use voice quality metrics of raw speech to assess dysarthria. Instead of using raw speech, a deep learning based diagnostic method that extracts these metrics from time-frequency speech representations will be reliable and preserves speaker identity. In this work a regression-based deep convolutional neural network is experimented to estimate jitter, shimmer, fundamental frequency (F0), and harmonic-to-noise ratio (HNR) employing 6 different time-frequency representations: spectrogram, low-frequency spectrogram, cepstrogram, low-frequency cepstrogram, cochleagram, and Mel scalogram. The experiment is assessed on VOC-ALS dysarthria speech dataset using RMSE and $$\text {R}^{2}$$ metrics, in which low-frequency cepstrogram performs well across all vowels, with average RMSEs of 0.76% for jitter, 2.36% for shimmer, and 4.075 dB for HNR whereas for F0, the cepstrogram shows the highest precision, with an average RMSE of 21.05 Hz. On Parkinson’s speech dataset (PC-GITA), the cepstrogram yields the lowest average RMSEs for jitter (0.57%), F0 (46.65 Hz), and HNR (3.80 dB), closely followed by the low-frequency cepstrogram with average RMSE for jitter (0.58%), F0 (48.53 Hz), and HNR (4.10 dB). For shimmer, the low-frequency cepstrogram attains the lowest average RMSE (2.64%). The results show that low-frequency cepstrogram and cepstrogram representation performs the best across all vowels for both dysarthric and Parkinson’s speech.
Authors
- Prakash Ramachandran (ORCID: https://orcid.org/0000-0001-5996-2413)
- Rajesh Kumar (ORCID: https://orcid.org/0000-0002-6019-0702)
- Aurobindo S
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1038/s41598-026-70228-8
- Primary Topic
- Voice and Speech Disorders
- Type
- article
- Field-Weighted Citation Impact
- 0.00