RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data

Public surgical segmentation datasets contain information about the operative view that extends beyond their pixel labels. RoboSurg-VQA organises this information as a closed-set visual question answering (VQA) benchmark while retaining the source of each answer. It contains 5632 frames from 29 EndoVis 2017/2018 source sequences and 11 task units, with answers derived from source metadata, mask measurements, or machine-generated candidate labels. Seven visual attributes were examined in a blinded 250-frame audit by two study-team reviewers working independently, with disagreements resolved by a third reviewer. We trained a shared model with frozen BiomedCLIP encoders and a 13-answer vocabulary to examine the contributions of image and question inputs. Against mask-derived and candidate labels on held-out sequences, the model achieved a seven-task mean fixed-label Macro-F1 of 0.540. Answer frequency and task-restricted image-only prediction scored 0.391 and 0.400, respectively. With unseen paraphrases, the score was 0.528. On 40 audited frames from held-out sequences, mean Macro-F1 across bleeding, image quality, and glare was 0.535 against the human reference, compared with 0.450 for answer frequency, a gain of 0.085 driven by bleeding. The audit identified attribute-specific label errors, with the greatest reviewer disagreement for glare.

Authors

Institutions

Publication Details

Journal
Machine Learning and Knowledge Extraction
Published
2026-09-28
DOI
https://doi.org/10.3390/make8100300
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data

Ziyang Wang, Chengyi Zhang, Zi Ye
Machine Learning and Knowledge Extraction
Multimodal Machine Learning Applications
article

RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data

Ziyang Wang, Chengyi Zhang, Zi Ye
article en

Abstract

Public surgical segmentation datasets contain information about the operative view that extends beyond their pixel labels. RoboSurg-VQA organises this information as a closed-set visual question answering (VQA) benchmark while retaining the source of each answer. It contains 5632 frames from 29 EndoVis 2017/2018 source sequences and 11 task units, with answers derived from source metadata, mask measurements, or machine-generated candidate labels. Seven visual attributes were examined in a blinded 250-frame audit by two study-team reviewers working independently, with disagreements resolved by a third reviewer. We trained a shared model with frozen BiomedCLIP encoders and a 13-answer vocabulary to examine the contributions of image and question inputs. Against mask-derived and candidate labels on held-out sequences, the model achieved a seven-task mean fixed-label Macro-F1 of 0.540. Answer frequency and task-restricted image-only prediction scored 0.391 and 0.400, respectively. With unseen paraphrases, the score was 0.528. On 40 audited frames from held-out sequences, mean Macro-F1 across bleeding, image quality, and glare was 0.535 against the human reference, compared with 0.450 for answer frequency, a gain of 0.085 driven by bleeding. The audit identified attribute-specific label errors, with the greatest reviewer disagreement for glare.

Machine Learning and Knowledge ExtractionVol. 8(10)
National University of Ireland, Maynooth (IE), Aston University (GB), Swansea University (GB)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

RoboSurg-VQA: A Multimodal Visual Question Answering Benchmark Derived from Surgical Segmentation Data — Ziyang Wang, Chengyi Zhang, et al. · Machine Learning and Knowledge Extraction (2026) | TGRS Research Map | TGRS