SurgCouple: Coupling Global Context with Local Cues via Chain of Spatio-temporal Evidence for Surgical Video Understanding
Abstract
Similar surgical phases and instrument configurations can involve distinct instrument–tissue interactions, so global context alone cannot discriminate local events. Existing multimodal LLMs, however, lack a mechanism for selectively attending to sparse, spatial-temporal cues. We introduce **SurgCouple**, the first surgical video reasoning framework to **couple global context with hypothesis-guided spatial-temporal evidence discovery**. An Evidence Chain progressively discovers temporal, spatial, and fine-grained evidence, with intermediate reasoning states providing hypothesis-guided feedback to the latter two stages to focus on the informative regions. A Reasoning Chain maintains an evolving interpretation of the surgical event, refining the hypothesis space as evidence accumulates. A Consistency Verifier ensures that local judgments remain grounded in the evidence actually selected and coherent across stages. With a frozen Qwen3-VL-8B, SurgCouple achieves 89.31% final-answer accuracy and a 58.92% valid chain rate, surpassing 81.94% and 34.01% for Qwen3-VL-8B, and 88.92% and 52.63% for GPT-5.5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.