Analyzing Dynamic Surgical Workflows Through Multi-Scale Vision-Language Reasoning
Abstract
Analyzing surgical workflow is critical for understanding complex procedural dynamics of a surgery. Current works focus on surgical workflow recognition to classify video streams into predetermined workflow sequences, inadequately representing the adaptive nature of clinical surgeries that respond to patient variability and evolving circumstances. We introduce the dynamic surgical workflow reasoning task, which eliminates fixed workflow constraints to address these limitations. Supporting this paradigm shift, we present DySurg (Dynamic Surgical Workflow), a comprehensive dataset containing over 100 hours of long surgical videos in real-world clinical recordings with annotated dynamic workflows across 7 major surgical categories. To enhance reasoning capabilities for this analysis, we propose an commentary-aligned video reasoning framework that constructs top-down visual reasoning sequences modeling surgeons' cognitive processes. Our approach aligns visual embeddings from surgical videos with semantic information extracted from video title and expert commentary, thereby aligning explainability with the visual representations. During inference, our model maps video frames to the semantic space to generate appropriate workflows, without the need of expert commentary. Extensive evaluations on the DySurg dataset demonstrates that our approach significantly outperforms existing large vision-language models (e.g. QWen2.5-VL) and surgical-specific pre-trained models (e.g., SurgVLP) in recognizing dynamic surgical workflows. All code, data, and models will be publicly released after the review process concludes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.