SurgAnt-ViVQA: learning to anticipate surgical events through GRU-driven temporal cross-attention

PURPOSE: Anticipating forthcoming surgical events is vital for real-time assistance in endonasal transsphenoidal pituitary surgery, where visibility is limited and workflow changes rapidly. Most visual question answering (VQA) systems reason on isolated frames with static vision-language alignment, providing limited support for forecasting forthcoming steps, instrument needs, or remaining procedure time. Existing surgical VQA datasets likewise focus primarily on the current scene rather than the near future. METHODS: We introduce PitVQA-Anticipation, a VQA dataset designed for forward-looking surgical reasoning. It comprises 33.5 h of operative video and 734,769 question-answer pairs built from temporally grouped clips and expert annotations across four tasks: predicting the future phase, next step, upcoming instrument, and remaining duration. We further propose SurgAnt-ViVQA, a video-language model that adapts a large language model using a GRU-Gated Temporal Cross-Attention module. A bidirectional GRU encodes frame-to-frame dynamics, while an adaptive gate injects visual context into the language stream at the token level. Parameter-efficient fine-tuning customizes the language backbone to the surgical domain. RESULTS: On PitVQA-Anticipation, SurgAnt-ViVQA achieved BLEU-4 72.38, ROUGE-L 84.94, and METEOR 87.05, outperforming the evaluated image-based and video-based baselines. It also generalized to the EndoVis18-VQA benchmark. The ablation results support the effectiveness of the combined GRU-gated temporal fusion module for anticipatory VQA, while further controlled experiments are required to fully disentangle the independent contributions of recurrence and adaptive gating. A frame-budget study indicates a trade-off: 8 frames maximize linguistic overlap metrics, whereas 32 frames slightly reduce BLEU but improve numeric time estimation. CONCLUSION: By pairing a temporally aware encoder with fine-grained gated cross-attention, SurgAnt-ViVQA advances surgical VQA from retrospective description toward proactive anticipation. PitVQA-Anticipation provides a computational benchmark for future-aware surgical reasoning; prospective clinical utility requires further validation with surgeons and operating-room staff.

Authors

Institutions

Publication Details

Journal
International Journal of Computer Assisted Radiology and Surgery
Published
2026-08-26
DOI
https://doi.org/10.1007/s11548-026-03784-z
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SurgAnt-ViVQA: learning to anticipate surgical events through GRU-driven temporal cross-attention

Matthew J. Clarkson, Danyal Z. Khan, Jiayuan Huang, Evangelos B. Mazomenos et al.
International Journal of Computer Assisted Radiology and Surgery
Multimodal Machine Learning Applications
article

SurgAnt-ViVQA: learning to anticipate surgical events through GRU-driven temporal cross-attention

Matthew J. Clarkson, Danyal Z. Khan, Jiayuan Huang, Evangelos B. Mazomenos, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Runlong He, Shreyas C. Dhake
article en

Abstract

PURPOSE: Anticipating forthcoming surgical events is vital for real-time assistance in endonasal transsphenoidal pituitary surgery, where visibility is limited and workflow changes rapidly. Most visual question answering (VQA) systems reason on isolated frames with static vision-language alignment, providing limited support for forecasting forthcoming steps, instrument needs, or remaining procedure time. Existing surgical VQA datasets likewise focus primarily on the current scene rather than the near future. METHODS: We introduce PitVQA-Anticipation, a VQA dataset designed for forward-looking surgical reasoning. It comprises 33.5 h of operative video and 734,769 question-answer pairs built from temporally grouped clips and expert annotations across four tasks: predicting the future phase, next step, upcoming instrument, and remaining duration. We further propose SurgAnt-ViVQA, a video-language model that adapts a large language model using a GRU-Gated Temporal Cross-Attention module. A bidirectional GRU encodes frame-to-frame dynamics, while an adaptive gate injects visual context into the language stream at the token level. Parameter-efficient fine-tuning customizes the language backbone to the surgical domain. RESULTS: On PitVQA-Anticipation, SurgAnt-ViVQA achieved BLEU-4 72.38, ROUGE-L 84.94, and METEOR 87.05, outperforming the evaluated image-based and video-based baselines. It also generalized to the EndoVis18-VQA benchmark. The ablation results support the effectiveness of the combined GRU-gated temporal fusion module for anticipatory VQA, while further controlled experiments are required to fully disentangle the independent contributions of recurrence and adaptive gating. A frame-budget study indicates a trade-off: 8 frames maximize linguistic overlap metrics, whereas 32 frames slightly reduce BLEU but improve numeric time estimation. CONCLUSION: By pairing a temporally aware encoder with fine-grained gated cross-attention, SurgAnt-ViVQA advances surgical VQA from retrospective description toward proactive anticipation. PitVQA-Anticipation provides a computational benchmark for future-aware surgical reasoning; prospective clinical utility requires further validation with surgeons and operating-room staff.

International Journal of Computer Assisted Radiology and Surgery
University of Manchester (GB), UCL Biomedical Research Centre (GB), The London College (GB), National Hospital for Neurology and Neurosurgery (GB), University College London (GB)
Cleveland Clinic, National Institute for Health and Care Research, University College London Hospitals NHS Foundation Trust, Engineering and Physical Sciences Research Council, UCLH Biomedical Research Centre
Openalex Percentile: Top 98%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.