Artificial Intelligence-Based Hypernasality Diagnosis Using CAPS-A-AM Rated Speech Samples in Pediatric Velopharyngeal Dysfunction
ObjectivePerceptual evaluation by speech-language pathologists (SLPs) is essential for initial evaluation of velopharyngeal dysfunction (VPD). Machine learning (ML) offers a promising avenue for developing accessible speech assessment tools when SLP expertise is limited. We aimed to develop a workflow and preliminary algorithm for AI-based hypernasality detection based on the Cleft Audit Protocol for Speech-Augmented-Americleft Modification (CAPS-A-AM), a standardized framework for auditory perceptual speech assessment.DesignIn this prospective, single-center study, speech samples were collected during SLP-guided evaluation, with consensus CAPS-A-AM ratings established. Mel spectrograms for high vowels (/i/ and /u/) were generated from sustained vowels, isolated words, and sentences for model development. Three ML approaches were evaluated: logistic regression (LR), Convolutional Neural Network (CNN) Attention-Multiple Instance Learning (MIL) (EfficientNet-V2-S), and a CNN-Extreme Gradient Boosting Hybrid (XGBoost).Patients/ParticipantsForty pediatric participants aged 2 to 17, including individuals with VPD, conditions associated with VPD, and healthy participants.Main Outcome Measure(s)Model performance was tested in binary hypernasality classification compared to SLP consensus at two CAPS-A-AM thresholds: absent (0) versus any hypernasality (1-4) and absent/borderline (0-1) versus mild-to-severe hypernasality (2-4).ResultsMultiple independent modeling approaches were able to detect clinically rated hypernasality. EfficientNet-V2-S achieved the highest observed performance estimates in this cohort (F1 score = 0.786-0.900). However, overlapping confidence intervals precluded demonstration of model superiority.ConclusionML-based hypernasality classification using mel spectrograms demonstrated promising feasibility. Future work will expand sample acquisition and refine model development with increasingly diverse speech inputs for training and sample rating.
Authors
- Molly F. MacIsaac (ORCID: https://orcid.org/0000-0002-3632-4205)
- Ruth Huntley Bahr (ORCID: https://orcid.org/0000-0002-7164-4685)
- Chelsea L. Sommer (ORCID: https://orcid.org/0000-0002-5552-6312)
- Luis Ahumada (ORCID: https://orcid.org/0000-0002-6856-6698)
- Jordan N. Halsey (ORCID: https://orcid.org/0000-0001-5513-5162)
- S. Alex Rottgers (ORCID: https://orcid.org/0000-0001-7209-1966)
- Joshua M. Wright (ORCID: https://orcid.org/0000-0003-2225-3247)
- Jamilla Vieux
- Mbinui N. Ghogomu (ORCID: https://orcid.org/0009-0009-5958-6193)
Institutions
- Johns Hopkins University (US)
- Florida International University (US)
- University of South Florida (US)
- Johns Hopkins All Children's Hospital (US)
Publication Details
- Journal
- The Cleft Palate-Craniofacial Journal
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1177/10556656261487525
- Primary Topic
- Cleft Lip and Palate Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00