Speaker-Independent Stuttering Detection Using Deep CNNs and Cepstral Acoustic Representations

Stuttering is a speech-fluency disorder whose clinical assessment has traditionally relied on manual perceptual judgment, a process that is time-consuming and susceptible to inter-rater variability. This paper presents a speaker-independent binary classification framework that distinguishes stuttered from non-stuttered speech using a convolutional neural network. Recordings from a normal-speech corpus and a stuttering corpus (originally labelled Normal, Fluent, and Dysfluent) were consolidated into two classes: Normal and Fluent were mapped to Non-stuttered, and Dysfluent to Stuttered, yielding 13,534 valid recordings after quality-control filtering. Each recording was resampled to 16 kHz mono, silence-trimmed, duration-filtered, and peak-amplitude normalised. Acoustic content was represented as 40-dimensional Mel-frequency cepstral coefficients with their first- and second-order temporal derivatives, giving a per-sample feature tensor of 120 × 313. Speaker identities were extracted and grouped before partitioning so that no speaker in the training set appeared in validation or test, explicitly preventing speaker leakage across 44 unique speakers (30 train / 7 validation / 7 test). On the held-out, speaker-disjoint test set of 2,152 recordings, the proposed CNN achieved 91.54% accuracy, 71.77% precision, 96.35% recall, an F1-score of 82.26%, and a ROC-AUC of 96.89% for the Stuttered class. All metrics were numerically re-derived from the confusion matrix and found internally consistent. The high recall on the minority Stuttered class indicates the model rarely misses true stuttering events, at the cost of more false positives — a trade-off discussed in the context of screening applications. The study remains limited by dataset scale, class imbalance, and the absence of cross-corpus and clinical validation. No claim of clinical or diagnostic validity is made.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22799577
Primary Topic
Stuttering Research and Treatment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Speaker-Independent Stuttering Detection Using Deep CNNs and Cepstral Acoustic Representations

Laiba Aamir, Fatima Maqbool Hussain
Zenodo (CERN European Organization for Nuclear Research)
Stuttering Research and Treatment
article

Speaker-Independent Stuttering Detection Using Deep CNNs and Cepstral Acoustic Representations

Laiba Aamir, Fatima Maqbool Hussain
article en

Abstract

Stuttering is a speech-fluency disorder whose clinical assessment has traditionally relied on manual perceptual judgment, a process that is time-consuming and susceptible to inter-rater variability. This paper presents a speaker-independent binary classification framework that distinguishes stuttered from non-stuttered speech using a convolutional neural network. Recordings from a normal-speech corpus and a stuttering corpus (originally labelled Normal, Fluent, and Dysfluent) were consolidated into two classes: Normal and Fluent were mapped to Non-stuttered, and Dysfluent to Stuttered, yielding 13,534 valid recordings after quality-control filtering. Each recording was resampled to 16 kHz mono, silence-trimmed, duration-filtered, and peak-amplitude normalised. Acoustic content was represented as 40-dimensional Mel-frequency cepstral coefficients with their first- and second-order temporal derivatives, giving a per-sample feature tensor of 120 × 313. Speaker identities were extracted and grouped before partitioning so that no speaker in the training set appeared in validation or test, explicitly preventing speaker leakage across 44 unique speakers (30 train / 7 validation / 7 test). On the held-out, speaker-disjoint test set of 2,152 recordings, the proposed CNN achieved 91.54% accuracy, 71.77% precision, 96.35% recall, an F1-score of 82.26%, and a ROC-AUC of 96.89% for the Stuttered class. All metrics were numerically re-derived from the confusion matrix and found internally consistent. The high recall on the minority Stuttered class indicates the model rarely misses true stuttering events, at the cost of more false positives — a trade-off discussed in the context of screening applications. The study remains limited by dataset scale, class imbalance, and the absence of cross-corpus and clinical validation. No claim of clinical or diagnostic validity is made.

Zenodo (CERN European Organization for Nuclear Research)
CytRx (United States) (US)
Quality Education
Openalex Percentile: Top 7%
Stuttering Research and Treatment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.