Towards Higher-Level Surgical Intelligence: A Comprehensive Benchmark and Multi-Agent Framework for Surgical Intent Understanding
Abstract
In recent years, surgical video understanding has progressed from instrument, action, and phase recognition to multimodal video question answering. However, to achieve higher-level surgical intelligence that can assist surgical maneuvers and support clinical decision-making, models need to understand the clinical intent underlying surgical actions: why an ongoing maneuver is performed and which subsequent surgical objective it prepares for. Therefore, we introduce the Surgical Intent Understanding task for endoscopic surgical videos, which jointly evaluates four clinical capabilities: anatomical localization, interpretation of instrument roles, interpretation of instrument coordination, and inference of subsequent surgical objectives. In consultation with experienced surgeons, we define four corresponding evaluation dimensions and construct the SurgIntent-VidQA dataset, comprising 603 video clips and 2,520 surgeon-annotated and reviewed five-option, single-answer questions across four procedure types in liver and lung surgery. Based on this dataset, we establish SurgIntent-Bench for systematic evaluation. We further propose SurgIntent-Agent, a knowledge-guided multi-agent framework that generates candidate intent hypotheses using surgical knowledge, acquires video evidence through targeted questions, and iteratively verifies and refines its predictions. We train the agents using supervised fine-tuning and direct preference optimization. Systematic evaluation on SurgIntent-Bench covers mainstream general-purpose, open-source video, medical, and task-specific fine-tuned multimodal large language models. SurgIntent-Agent achieves an overall accuracy of 67.62%, outperforming the strongest general-purpose model (61.32%) and task-specific SFT baseline (63.72%) by 6.30 and 3.90 percentage points, respectively. Ablation studies further support the contributions of surgical knowledge, visual evidence extraction, and preference optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.