HeiCo-FOCUS: A Clinically Grounded Dataset for Long-Context Video Understanding
Abstract
Recent advances in Vision-Language Models (VLMs) have led to rapid progress in video understanding across a wide range of benchmark tasks. However, ex- isting evaluations largely focus on short-term reasoning, failing to assess a criti- cal capability: maintaining cumulative temporal consistency over extended time horizons. To close this evaluation gap, we introduce HeiCo-FOCUS, a clinically grounded dataset for evaluating long-context video understanding through the task of **F**oreign **O**bject **C**ontextual **U**nderstanding in **S**urgery. Built on a dataset of **He**idelberg **Co**lorectal surgeries, this task requires models to continuously track multiple objects as they are inserted, manipulated, occluded, and removed over procedures lasting up to hours. HeiCo-FOCUS comprises 30,000 visual question answerings (VQAs) pairs covering five core capabilities: object recognition, tem- poral grounding, aggregation, event and procedural understanding, and complex reasoning. The dataset was constructed through a rigorous multi-stage annotation pipeline involving large-scale crowd annotation and 39 surgical domain experts to ensure high quality and clinical relevance. To systematically probe model behav- ior, we introduce a multi-track evaluation framework that progressively increases temporal and contextual demands from single frames to full procedures. Experi- ments with ten frontier VLMs show that HeiCo-FOCUS tasks are far from solved: only around half of the models clearly outperform a text-only baseline. Across the video tracks, models perform best on event and procedural understanding (mean Accuracy: 56.5% across all models), while temporal grounding remains particu- larly challenging for all evaluated models (mean Accuracy: 19.7%). We therefore expect HeiCo-FOCUS to serve as a catalyst for the development of models capa- ble of reliable, temporally consistent reasoning over hours-long videos.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.