LUCID: Intelligent Informative Frame Selection in Otoscopy for Enhanced Diagnostic Utility

Background/Objectives: Accurate diagnosis of middle ear disease remains challenging because otoscopic interpretation is subjective, while many artificial intelligence models depend on manually selected still frames from video examinations. We aimed to develop and internally evaluate LUCID, a systematic framework for automatically identifying the most informative frames in otoscopy videos. Methods: We analyzed 713 otoscopy videos from 491 patients. LUCID combines a ResNet-50 classifier for coarse eardrum-visibility assessment, BC-AdvCAM for weakly supervised eardrum localization and coverage estimation, and an otoscope-specific blur/focus assessment into a within-video frame-ranking score. Automated selections were compared with expert-selected frames through blinded human review and downstream classification. To separate ranking quality from the effects of bag size and multi-frame aggregation, a matched 12-frame control using uniform temporal sampling was evaluated with the same ResNet-50 and attention-based multiple instance learning (ABMIL) pipeline. Results: Among 293 video pairs with complete ratings, algorithm-selected frames were rated equal to or better than expert-selected frames in 74.7% of evaluations on average. However, when an evaluator expressed a non-Equal preference, the expert frame was favored in 71.7% and 70.7% of cases by the two evaluators, respectively; among concordant non-Equal ratings, 80.0% favored the expert frame. BC-AdvCAM achieved an intersection over union of 0.6717 and a Dice similarity coefficient of 0.7924 on a held-out manually segmented test subset. For downstream diagnosis, patient-level mean accuracy for the matched single-frame comparison was 68.79% for LUCID Top-1 versus 72.17% for the expert-selected frame (difference, −3.38 percentage points; 95% exact sign-flip-inversion CI, −6.79 to 0.00; two-sided exact patient sign-flip p = 0.0540). In the matched 12-frame comparison, LUCID Top-12 achieved 78.13% versus 59.41% for Uniform-12 using the identical ABMIL pipeline (difference, +18.71 percentage points; 95% exact sign-flip-inversion CI, +15.79 to +21.63; two-sided exact patient sign-flip p = 2.33 × 10−10). Conclusions: In the matched single-frame comparison, the expert-selected frame had the higher observed patient-level mean accuracy, whereas LUCID provided an automated ranking of the full video that substantially improved matched 12-frame ABMIL performance relative to uniform temporal sampling. These results support LUCID as an automated frame-ranking and multi-frame evidence-selection tool, while external validation remains necessary before clinical deployment.

Authors

Institutions

Publication Details

Journal
Diagnostics
Published
2026-09-15
DOI
https://doi.org/10.3390/diagnostics16182981
Primary Topic
Ear Surgery and Otitis Media
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

LUCID: Intelligent Informative Frame Selection in Otoscopy for Enhanced Diagnostic Utility

Aaron C. Moberly, Hao Lü, Muhammad Khalid Khan Niazi, Metin N. Gürcan et al.
Diagnostics
Ear Surgery and Otitis Media
article

LUCID: Intelligent Informative Frame Selection in Otoscopy for Enhanced Diagnostic Utility

Aaron C. Moberly, Hao Lü, Muhammad Khalid Khan Niazi, Metin N. Gürcan, Delal Şeker, Amy Zinnia, Memnun Demir, Gabriella I. Puchall, Tucker Corwen, Carl D. Langefeld, Shalaka Chavan, Zian Shang
article en

Abstract

Background/Objectives: Accurate diagnosis of middle ear disease remains challenging because otoscopic interpretation is subjective, while many artificial intelligence models depend on manually selected still frames from video examinations. We aimed to develop and internally evaluate LUCID, a systematic framework for automatically identifying the most informative frames in otoscopy videos. Methods: We analyzed 713 otoscopy videos from 491 patients. LUCID combines a ResNet-50 classifier for coarse eardrum-visibility assessment, BC-AdvCAM for weakly supervised eardrum localization and coverage estimation, and an otoscope-specific blur/focus assessment into a within-video frame-ranking score. Automated selections were compared with expert-selected frames through blinded human review and downstream classification. To separate ranking quality from the effects of bag size and multi-frame aggregation, a matched 12-frame control using uniform temporal sampling was evaluated with the same ResNet-50 and attention-based multiple instance learning (ABMIL) pipeline. Results: Among 293 video pairs with complete ratings, algorithm-selected frames were rated equal to or better than expert-selected frames in 74.7% of evaluations on average. However, when an evaluator expressed a non-Equal preference, the expert frame was favored in 71.7% and 70.7% of cases by the two evaluators, respectively; among concordant non-Equal ratings, 80.0% favored the expert frame. BC-AdvCAM achieved an intersection over union of 0.6717 and a Dice similarity coefficient of 0.7924 on a held-out manually segmented test subset. For downstream diagnosis, patient-level mean accuracy for the matched single-frame comparison was 68.79% for LUCID Top-1 versus 72.17% for the expert-selected frame (difference, −3.38 percentage points; 95% exact sign-flip-inversion CI, −6.79 to 0.00; two-sided exact patient sign-flip p = 0.0540). In the matched 12-frame comparison, LUCID Top-12 achieved 78.13% versus 59.41% for Uniform-12 using the identical ABMIL pipeline (difference, +18.71 percentage points; 95% exact sign-flip-inversion CI, +15.79 to +21.63; two-sided exact patient sign-flip p = 2.33 × 10−10). Conclusions: In the matched single-frame comparison, the expert-selected frame had the higher observed patient-level mean accuracy, whereas LUCID provided an automated ranking of the full video that substantially improved matched 12-frame ABMIL performance relative to uniform temporal sampling. These results support LUCID as an automated frame-ranking and multi-frame evidence-selection tool, while external validation remains necessary before clinical deployment.

DiagnosticsVol. 16(18)
Dicle University (TR), Wake Forest University (US), The Ohio State University (US), Vanderbilt University Medical Center (US)
National Institute on Deafness and Other Communication Disorders
Quality Education
Openalex Percentile: Top 9%
Ear Surgery and Otitis Media
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.