Real‐Time Classification of Functional Laryngeal Behaviors From Vocal Fold Kinematics

{"OBJECTIVES:":[0],"Develop":[1],"a":[2,41,132,230],"real-time":[3],"prediction":[4],"model":[5,48,101,130,168],"to":[6,27,65,187],"detect":[7],"laryngeal":[8,56,69,160,205,240],"behaviors":[9],"from":[10,59,95,183,207],"videolaryngoscopy-based":[11,208],"tracking":[12,52,210],"data":[13,20,53],"during":[14,211],"clinical":[15,212],"examinations.":[16],"This":[17],"will":[18],"improve":[19],"collection":[21],"quality":[22],"by":[23],"providing":[24],"immediate":[25],"feedback":[26],"clinicians":[28],"and":[29,81,113,122,138,155,225,237],"enable":[30],"future":[31],"automation":[32],"pipelines":[33],"for":[34,185,189,197,232],"efficient":[35],"large-scale":[36],"analysis.":[37],"METHODS:":[38],"We":[39],"trained":[40],"stateful":[42],"residual":[43],"Gated":[44],"Recurrent":[45],"Unit":[46],"(GRU)":[47],"that":[49],"analyzed":[50],"the":[51,67,146,164,167,179],"of":[54,83,136,141,159,204,220,239],"39":[55],"keypoints,":[57],"derived":[58],"our":[60],"published":[61],"keypoint":[62],"detection":[63,219],"model,":[64],"predict":[66],"patient's":[68],"task":[70,161],"state.":[71],"These":[72],"included":[73],"\\"phonation,\\"":[74],"\\"sustained":[75],"phonation,\\"":[76],"\\"swallowing,\\"":[77],"\\"idle,\\"":[78],"\\"coughing,\\"":[79],"\\"sniffing,\\"":[80],"\\"out":[82],"view.\\"":[84],"Model":[85],"development":[86],"used":[87],"916":[88],"state":[89,173],"segments":[90],"comprising":[91,108],"222,065":[92],"video":[93],"frames":[94],"72":[96],"laryngoscopy":[97],"videos.":[98],"The":[99,129,191],"resulting":[100],"was":[102,117,195],"evaluated":[103,144],"on":[104,145,178],"an":[105,139],"independent":[106,147],"dataset":[107],"8":[109],"videos,":[110],"123":[111],"segments,":[112],"49,770":[114],"frames.":[115],"Performance":[116],"assessed":[118],"using":[119],"classification":[120,203],"metrics":[121],"temporal":[123,157],"intersection":[124],"over":[125],"union":[126],"(mIoU).":[127],"RESULTS:":[128],"achieved":[131],"mean":[133],"accuracy":[134],"score":[135],"92%":[137],"mIoU":[140],"0.82":[142],"when":[143],"test":[148,180],"dataset,":[149],"indicating":[150],"agreement":[151],"with":[152,175],"manual":[153],"annotations":[154],"accurate":[156],"identification":[158],"states.":[162],"At":[163],"class":[165],"level,":[166],"performed":[169],"consistently":[170],"across":[171],"most":[172],"categories,":[174],"per-class":[176],"F1-scores":[177],"set":[181],"ranging":[182],"83%":[184],"phonation":[186],"97%":[188],"sniffing.":[190],"lowest":[192],"validation":[193],"F1":[194],"observed":[196],"cough":[198],"(82%":[199],"[70%-92%]).":[200],"CONCLUSION:":[201],"Real-time":[202],"states":[206,221],"pose":[209],"examinations":[213],"is":[214],"feasible.":[215],"With":[216],"reliable":[217],"automated":[218,234],"such":[222],"as":[223],"swallowing":[224],"phonation,":[226],"this":[227],"work":[228],"establishes":[229],"foundation":[231],"further":[233],"analysis,":[235],"interpretation,":[236],"documentation":[238],"pathology.":[241]}

Authors

Institutions

Publication Details

Journal
The Laryngoscope
Published
2026-09-18
DOI
https://doi.org/10.1002/lary.70927
Primary Topic
Voice and Speech Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Real‐Time Classification of Functional Laryngeal Behaviors From Vocal Fold Kinematics

Kristina Simonyan, Matthew R. Naunheim, Aki Koivu, Pin‐Yu Lin
The Laryngoscope
Voice and Speech Disorders
article

Real‐Time Classification of Functional Laryngeal Behaviors From Vocal Fold Kinematics

Kristina Simonyan, Matthew R. Naunheim, Aki Koivu, Pin‐Yu Lin
article en

Abstract

OBJECTIVES: Develop a real-time prediction model to detect laryngeal behaviors from videolaryngoscopy-based tracking data during clinical examinations. This will improve data collection quality by providing immediate feedback to clinicians and enable future automation pipelines for efficient large-scale analysis. METHODS: We trained a stateful residual Gated Recurrent Unit (GRU) model that analyzed the tracking data of 39 laryngeal keypoints, derived from our published keypoint detection model, to predict the patient's laryngeal task state. These included "phonation," "sustained phonation," "swallowing," "idle," "coughing," "sniffing," and "out of view." Model development used 916 state segments comprising 222,065 video frames from 72 laryngoscopy videos. The resulting model was evaluated on an independent dataset comprising 8 videos, 123 segments, and 49,770 frames. Performance was assessed using classification metrics and temporal intersection over union (mIoU). RESULTS: The model achieved a mean accuracy score of 92% and an mIoU of 0.82 when evaluated on the independent test dataset, indicating agreement with manual annotations and accurate temporal identification of laryngeal task states. At the class level, the model performed consistently across most state categories, with per-class F1-scores on the test set ranging from 83% for phonation to 97% for sniffing. The lowest validation F1 was observed for cough (82% [70%-92%]). CONCLUSION: Real-time classification of laryngeal states from videolaryngoscopy-based pose tracking during clinical examinations is feasible. With reliable automated detection of states such as swallowing and phonation, this work establishes a foundation for further automated analysis, interpretation, and documentation of laryngeal pathology.

The Laryngoscope
Massachusetts Eye and Ear Infirmary (US), Harvard University (US)
Openalex Percentile: Top 11%
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.