Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network

Abstract Speech-language pathologists use voice quality metrics of raw speech to assess dysarthria. Instead of using raw speech, a deep learning based diagnostic method that extracts these metrics from time-frequency speech representations will be reliable and preserves speaker identity. In this work a regression-based deep convolutional neural network is experimented to estimate jitter, shimmer, fundamental frequency (F0), and harmonic-to-noise ratio (HNR) employing 6 different time-frequency representations: spectrogram, low-frequency spectrogram, cepstrogram, low-frequency cepstrogram, cochleagram, and Mel scalogram. The experiment is assessed on VOC-ALS dysarthria speech dataset using RMSE and $$\text {R}^{2}$$ metrics, in which low-frequency cepstrogram performs well across all vowels, with average RMSEs of 0.76% for jitter, 2.36% for shimmer, and 4.075 dB for HNR whereas for F0, the cepstrogram shows the highest precision, with an average RMSE of 21.05 Hz. On Parkinson’s speech dataset (PC-GITA), the cepstrogram yields the lowest average RMSEs for jitter (0.57%), F0 (46.65 Hz), and HNR (3.80 dB), closely followed by the low-frequency cepstrogram with average RMSE for jitter (0.58%), F0 (48.53 Hz), and HNR (4.10 dB). For shimmer, the low-frequency cepstrogram attains the lowest average RMSE (2.64%). The results show that low-frequency cepstrogram and cepstrogram representation performs the best across all vowels for both dysarthric and Parkinson’s speech.

Authors

Publication Details

Journal
Scientific Reports
Published
2026-09-29
DOI
https://doi.org/10.1038/s41598-026-70228-8
Primary Topic
Voice and Speech Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network

Prakash Ramachandran, Rajesh Kumar, Aurobindo S
Scientific Reports
Voice and Speech Disorders
article

Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network

Prakash Ramachandran, Rajesh Kumar, Aurobindo S
article en

Abstract

Abstract Speech-language pathologists use voice quality metrics of raw speech to assess dysarthria. Instead of using raw speech, a deep learning based diagnostic method that extracts these metrics from time-frequency speech representations will be reliable and preserves speaker identity. In this work a regression-based deep convolutional neural network is experimented to estimate jitter, shimmer, fundamental frequency (F0), and harmonic-to-noise ratio (HNR) employing 6 different time-frequency representations: spectrogram, low-frequency spectrogram, cepstrogram, low-frequency cepstrogram, cochleagram, and Mel scalogram. The experiment is assessed on VOC-ALS dysarthria speech dataset using RMSE and $$\text {R}^{2}$$ metrics, in which low-frequency cepstrogram performs well across all vowels, with average RMSEs of 0.76% for jitter, 2.36% for shimmer, and 4.075 dB for HNR whereas for F0, the cepstrogram shows the highest precision, with an average RMSE of 21.05 Hz. On Parkinson’s speech dataset (PC-GITA), the cepstrogram yields the lowest average RMSEs for jitter (0.57%), F0 (46.65 Hz), and HNR (3.80 dB), closely followed by the low-frequency cepstrogram with average RMSE for jitter (0.58%), F0 (48.53 Hz), and HNR (4.10 dB). For shimmer, the low-frequency cepstrogram attains the lowest average RMSE (2.64%). The results show that low-frequency cepstrogram and cepstrogram representation performs the best across all vowels for both dysarthric and Parkinson’s speech.

Scientific Reports
Openalex Percentile: Top 12%
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Estimation of voice quality parameters of dysarthric speech with preserved identity using different time-frequency image representations in deep learning network — Prakash Ramachandran, Rajesh Kumar, et al. · Scientific Reports (2026) | TGRS Research Map | TGRS