acceptodds
Under review as a conference paper at ICLR 2027

GATE: Can Vision-Language Models Ground, Audit, and Track Surgical Events?

Abstract

Vision-language models (VLMs) could one day assist surgeons, but only if they can reliably interpret visual evidence, verify their own interpretations, and follow procedures as they unfold. How well current VLMs support these abilities, and how training can improve them, remain open questions. We introduce GATE, an evaluation framework that integrates a surgical video benchmark with systematic training interventions to assess whether models can Ground, Audit, and Track surgical Events. We therefore establish a capability taxonomy of six categories and 15 tasks, supported by a benchmark of 13,715 human-reviewed questions over 36,264 video clips from eight datasets across six anatomical sites. We evaluate 11 leading proprietary and open-weight models, including the recently released GPT-6 Astra, which is the strongest model in our evaluation yet achieves only 49.0% overall accuracy. On average, models are stronger at recognition and spatial judgments than at workflow reasoning and joint description and workflow correction. This uneven profile makes procedural reasoning and consistent verification central priorities for advancing surgical intelligence. Training interventions offer a concrete path forward: across models adapted separately for each task from a backbone with two billion parameters, supervised fine-tuning followed by reinforcement learning raises pooled test accuracy from 2.3% to 38.9%, compared with 35.7% for supervised fine-tuning and 12.3% for direct reinforcement learning. Supervised fine-tuning accounts for most of the overall accuracy gain, with further improvement from subsequent reinforcement learning. GATE thus provides an assessment framework that quantifies current capability gaps, evaluates how training interventions address these gaps, and guides future research and model development in surgical intelligence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.