LUCID: Intelligent Informative Frame Selection in Otoscopy for Enhanced Diagnostic Utility
Background/Objectives: Accurate diagnosis of middle ear disease remains challenging because otoscopic interpretation is subjective, while many artificial intelligence models depend on manually selected still frames from video examinations. We aimed to develop and internally evaluate LUCID, a systematic framework for automatically identifying the most informative frames in otoscopy videos. Methods: We analyzed 713 otoscopy videos from 491 patients. LUCID combines a ResNet-50 classifier for coarse eardrum-visibility assessment, BC-AdvCAM for weakly supervised eardrum localization and coverage estimation, and an otoscope-specific blur/focus assessment into a within-video frame-ranking score. Automated selections were compared with expert-selected frames through blinded human review and downstream classification. To separate ranking quality from the effects of bag size and multi-frame aggregation, a matched 12-frame control using uniform temporal sampling was evaluated with the same ResNet-50 and attention-based multiple instance learning (ABMIL) pipeline. Results: Among 293 video pairs with complete ratings, algorithm-selected frames were rated equal to or better than expert-selected frames in 74.7% of evaluations on average. However, when an evaluator expressed a non-Equal preference, the expert frame was favored in 71.7% and 70.7% of cases by the two evaluators, respectively; among concordant non-Equal ratings, 80.0% favored the expert frame. BC-AdvCAM achieved an intersection over union of 0.6717 and a Dice similarity coefficient of 0.7924 on a held-out manually segmented test subset. For downstream diagnosis, patient-level mean accuracy for the matched single-frame comparison was 68.79% for LUCID Top-1 versus 72.17% for the expert-selected frame (difference, −3.38 percentage points; 95% exact sign-flip-inversion CI, −6.79 to 0.00; two-sided exact patient sign-flip p = 0.0540). In the matched 12-frame comparison, LUCID Top-12 achieved 78.13% versus 59.41% for Uniform-12 using the identical ABMIL pipeline (difference, +18.71 percentage points; 95% exact sign-flip-inversion CI, +15.79 to +21.63; two-sided exact patient sign-flip p = 2.33 × 10−10). Conclusions: In the matched single-frame comparison, the expert-selected frame had the higher observed patient-level mean accuracy, whereas LUCID provided an automated ranking of the full video that substantially improved matched 12-frame ABMIL performance relative to uniform temporal sampling. These results support LUCID as an automated frame-ranking and multi-frame evidence-selection tool, while external validation remains necessary before clinical deployment.
Authors
- Aaron C. Moberly (ORCID: https://orcid.org/0000-0001-9022-6916)
- Hao Lü (ORCID: https://orcid.org/0000-0002-4547-7780)
- Muhammad Khalid Khan Niazi (ORCID: https://orcid.org/0000-0002-1278-7512)
- Metin N. Gürcan (ORCID: https://orcid.org/0000-0002-2421-8229)
- Delal Şeker (ORCID: https://orcid.org/0000-0002-6863-7150)
- Amy Zinnia
- Memnun Demir (ORCID: https://orcid.org/0009-0000-3523-6986)
- Gabriella I. Puchall
- Tucker Corwen
- Carl D. Langefeld
- Shalaka Chavan
- Zian Shang
Institutions
- Dicle University (TR)
- Wake Forest University (US)
- The Ohio State University (US)
- Vanderbilt University Medical Center (US)
Publication Details
- Journal
- Diagnostics
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/diagnostics16182981
- Primary Topic
- Ear Surgery and Otitis Media
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Institute on Deafness and Other Communication Disorders