Machine Learning for Non‐invasive Differentiation of Glottic Carcinoma Versus Laryngeal Immobility

OBJECTIVE: Voice analysis for medical diagnostics is advancing rapidly with artificial intelligence and machine learning. Non-invasive voice analysis could optimize diagnostic procedures in ear, nose, and throat (ENT), supplementing endoscopic evaluations and guiding treatment strategies. This study aims to distinguish early glottic carcinomas (EGCs) from unilateral laryngeal immobility (ULI) in dysphonic patients using machine learning. STUDY DESIGN: A machine learning model was developed in Python, incorporating data segmentation, normalization, hyperparameter tuning, and data augmentation via parameter transformation and using Synthetic Minority Over-sampling Technique (SMOTE) to balance class distributions. A metamodel combining multiple algorithms (Random Forest Classifier, Support Vector Machines, and Extreme Gradient Boosting) enhanced predictive accuracy. SETTING: Tertiary Medical Center (laryngology unit, ENT Department). METHODS: Population is dysphonic patients diagnosed with ULI or EGC, and asymptomatic controls. Patients with other voice-affecting conditions were excluded. Sixty-three pathologic cases (ULI: 38, EGC: 25) and 40 controls. Data augmentation and rebalance expanded the data set to 650 samples. Voice recordings of sustained vowels were analyzed for vocal parameters (fundamental frequency, jitter, and shimmer) using machine learning. High-quality audio equipment was used. Model performance was evaluated by precision, recall, and area under receiver operating characteristic curve (AUC) in differentiating EGC from ULI. RESULTS: The model achieved 72% accuracy in distinguishing ULI from EGC, validated through tests. CONCLUSION: Machine learning shows strong potential for distinguishing EGC from ULI. This non-invasive method could improve diagnostic accuracy and enable earlier intervention, especially in regions with limited specialist access. Further refinement could enhance its role in public health.

Authors

Institutions

Publication Details

Journal
Otolaryngology
Published
2026-09-29
DOI
https://doi.org/10.1002/ohn.70465
Primary Topic
Voice and Speech Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning for Non‐invasive Differentiation of Glottic Carcinoma Versus Laryngeal Immobility

Robin Baudouin, Tiffany Rigal, Aude Julien‐Laferriere, Grégoire Vialatte de Pémille et al.
Otolaryngology
Voice and Speech Disorders
article

Machine Learning for Non‐invasive Differentiation of Glottic Carcinoma Versus Laryngeal Immobility

Robin Baudouin, Tiffany Rigal, Aude Julien‐Laferriere, Grégoire Vialatte de Pémille, Marta P. Circiu, Jerome Rene Lechien, Clémence Forges, Stéphane Hans, Lise Crevier‐Buchman, Stanislas Nicolleau
article en

Abstract

OBJECTIVE: Voice analysis for medical diagnostics is advancing rapidly with artificial intelligence and machine learning. Non-invasive voice analysis could optimize diagnostic procedures in ear, nose, and throat (ENT), supplementing endoscopic evaluations and guiding treatment strategies. This study aims to distinguish early glottic carcinomas (EGCs) from unilateral laryngeal immobility (ULI) in dysphonic patients using machine learning. STUDY DESIGN: A machine learning model was developed in Python, incorporating data segmentation, normalization, hyperparameter tuning, and data augmentation via parameter transformation and using Synthetic Minority Over-sampling Technique (SMOTE) to balance class distributions. A metamodel combining multiple algorithms (Random Forest Classifier, Support Vector Machines, and Extreme Gradient Boosting) enhanced predictive accuracy. SETTING: Tertiary Medical Center (laryngology unit, ENT Department). METHODS: Population is dysphonic patients diagnosed with ULI or EGC, and asymptomatic controls. Patients with other voice-affecting conditions were excluded. Sixty-three pathologic cases (ULI: 38, EGC: 25) and 40 controls. Data augmentation and rebalance expanded the data set to 650 samples. Voice recordings of sustained vowels were analyzed for vocal parameters (fundamental frequency, jitter, and shimmer) using machine learning. High-quality audio equipment was used. Model performance was evaluated by precision, recall, and area under receiver operating characteristic curve (AUC) in differentiating EGC from ULI. RESULTS: The model achieved 72% accuracy in distinguishing ULI from EGC, validated through tests. CONCLUSION: Machine learning shows strong potential for distinguishing EGC from ULI. This non-invasive method could improve diagnostic accuracy and enable earlier intervention, especially in regions with limited specialist access. Further refinement could enhance its role in public health.

Otolaryngology
Centre National de la Recherche Scientifique (FR), Université Sorbonne Nouvelle (FR), Université de Versailles Saint-Quentin-en-Yvelines (FR), Université Paris-Saclay (FR), Sorbonne Université (FR), ESME - École d’ingénieurs pluridisciplinaires (FR), Laboratoire de Phonétique et Phonologie (FR), Hôpital Foch (FR)
Openalex Percentile: Top 12%
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.