SurgVSI: A Benchmark for 3D Spatial Reasoning in Surgical Environments
Abstract
Robust 3D spatial reasoning is essential for understanding complex surgical scenes, where accurate perception of geometric relationships, depth, and motion is critical for downstream decision-making. However, existing benchmarks largely focus on general-domain settings and do not capture the fine-grained geometric interactions and dynamic constraints present in real surgical environments. In this work, we introduce SurgVSI, a benchmark dataset for evaluating 3D spatial reasoning in surgical environments. SurgVSI is constructed from StereoMIS, EndoVis17, and EndoVis18, and comprises a diverse set of tasks that assess both static and dynamic spatial understanding, including depth estimation, distance estimation, camera ego-motion analysis, and instrument path tracking and completion. The dataset contains 5,998 evaluation instances and 147,386 training instances, with metric annotations derived from stereo reconstruction and camera calibration. We conduct extensive evaluations across representative models to characterize current capabilities on SurgVSI, revealing substantial challenges in accurate geometric reasoning and dynamic spatial understanding. Our benchmark establishes a standardized testbed for studying 3D spatial reasoning in surgical contexts and provides a foundation for future research in perception, reasoning, and robotic systems operating in complex real-world environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.