Video-to-report generation for cataract surgery using procedurally grounded vision–language models
Cataract surgery is widely performed, but operative reports are usually written retrospectively and may omit short events or use inconsistent descriptions. We aimed to develop a system that converts unedited cataract surgery videos into temporally structured operative reports with inspectable visual evidence. We developed LensNarrate, which combines fixed-rate video sampling, vision-language captioning, temporal assembly of frame-level predictions, and separate visual grounding. We evaluated 20 Polish/Ukrainian videos using five-fold video-level cross-validation and 10 held-out Ningbo videos after target-domain adaptation. Performance was summarized at the video level with bootstrap 95% confidence intervals. Here we show that LensNarrate achieves internal temporal frame accuracy of 83.2% (95% confidence interval, 80.1–85.5), temporal macro-F1 of 66.9% (64.2–69.2), and segmental F1@50 of 58.0% (50.5–66.1). On the Ningbo held-out set, the corresponding results are 71.6% (66.4-76.2), 57.5% (52.5–62.1), and 47.7% (40.2–56.4). Performance is comparable to a fine-tuned Qwen2.5 vision-language model and higher than a fine-tuned VideoCLIP retrieval baseline. Temporal assembly increases internal segmental F1@50 from 6.5% to 58.0% and reduces over-segmentation from 14.4 to 1.56. LensNarrate provides a traceable workflow for producing structured cataract surgery reports from video. The results support further multi-center and prospective evaluation while identifying cross-site variation and short-phase boundary localization as remaining challenges. Cataract surgery removes the eye’s cloudy natural lens and replaces it with an artificial intraocular lens to improve vision. These operations are often recorded as videos, but the written operation report is usually prepared afterwards from memory. Important short events can therefore be missed or described inconsistently. We developed LensNarrate, a computer system that samples frames from a full surgery video, describes what is happening, arranges the descriptions in time order, and shows visual evidence for instruments and eye structures. We tested the system on videos from hospitals in Europe and China and compared it with two other computer methods. LensNarrate produced readable, phase-based reports and kept the main order of the operation, although performance was lower when videos came from a different hospital. Larger studies are needed before clinical use, but this approach may support more consistent documentation, review, and surgical training.
Authors
- Andrzej Grzybowski (ORCID: https://orcid.org/0000-0002-3724-2391)
- Kai Jin (ORCID: https://orcid.org/0000-0003-4369-2417)
- Andrii Ruban
- Kaikai Zhao
- Tao Yu
- Rui Yao
- Quanyong Yi
- Vitalii Prudyus
Institutions
- China University of Mining and Technology (CN)
- Wenzhou Medical University (CN)
- Affiliated Eye Hospital of Wenzhou Medical College (CN)
- National Academy of Medical Sciences of Ukraine (UA)
- Kyiv City Clinical Oncology Center (UA)
- Second Affiliated Hospital of Zhejiang University (CN)
- Kundiiev Institute of Occupational Health of the National Academy of Medical Sciences of Ukraine (UA)
- University of Warmia and Mazury in Olsztyn (PL)
Publication Details
- Journal
- Communications Medicine
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1038/s43856-026-01904-z
- Primary Topic
- Digital Imaging in Medicine
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Natural Science Foundation of China