acceptodds
Under review as a conference paper at ICLR 2027

SurgHiCo: Hierarchical Composition of Long-Horizon Surgical Video–Language Representations

Abstract

Surgical video-language pretraining uses narrated videos and their hierarchical captions, from individual maneuvers to stretches of a procedure. However, clip-based baselines encode segments with a fixed frame budget or aggregate local features by mean pooling, providing limited modeling of how local events form longer sequences. Under matched controls, neither incorporating hierarchical captions nor enlarging the temporal receptive field substantially improves the baseline, suggesting that the bottleneck lies in how longer-range representations are formed rather than in supervision or input. Motivated by this finding, we propose SurgHiCo, a hierarchical video-language pretraining framework that organizes representation learning by temporal scale, following the structure of surgery itself: maneuvers make up steps, and steps make up phases. Each level is composed from the level below by a lightweight composer and aligned with text at its own span. SurgHiCo achieves state-of-the-art zero-shot surgical phase recognition across five datasets and the strongest linear-probing performance at every labeled-video fraction on all of them. Notably, it improves average zero-shot accuracy by 6.4 points over the strongest prior model and full-shot linear-probing macro-F1 by 12.2 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.