RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data
Public surgical segmentation datasets contain information about the operative view that extends beyond their pixel labels. RoboSurg-VQA organises this information as a closed-set visual question answering (VQA) benchmark while retaining the source of each answer. It contains 5632 frames from 29 EndoVis 2017/2018 source sequences and 11 task units, with answers derived from source metadata, mask measurements, or machine-generated candidate labels. Seven visual attributes were examined in a blinded 250-frame audit by two study-team reviewers working independently, with disagreements resolved by a third reviewer. We trained a shared model with frozen BiomedCLIP encoders and a 13-answer vocabulary to examine the contributions of image and question inputs. Against mask-derived and candidate labels on held-out sequences, the model achieved a seven-task mean fixed-label Macro-F1 of 0.540. Answer frequency and task-restricted image-only prediction scored 0.391 and 0.400, respectively. With unseen paraphrases, the score was 0.528. On 40 audited frames from held-out sequences, mean Macro-F1 across bleeding, image quality, and glare was 0.535 against the human reference, compared with 0.450 for answer frequency, a gain of 0.085 driven by bleeding. The audit identified attribute-specific label errors, with the greatest reviewer disagreement for glare.
Authors
- Ziyang Wang (ORCID: https://orcid.org/0000-0003-1605-0873)
- Chengyi Zhang (ORCID: https://orcid.org/0009-0001-1966-0776)
- Zi Ye (ORCID: https://orcid.org/0000-0003-1002-0315)
Institutions
- National University of Ireland, Maynooth (IE)
- Aston University (GB)
- Swansea University (GB)
Publication Details
- Journal
- Machine Learning and Knowledge Extraction
- Published
- 2026-09-28
- DOI
- https://doi.org/10.3390/make8100300
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00