Controlled Transfer and Sparse Expert Routing for Surgical Foundation Models
Abstract
Adapting surgical workflow recognition to a new procedure requires choosing a visual representation with limited labeled videos. Existing comparisons often change the encoder, temporal model, and training recipe together, leaving it unclear which representations transfer well and when source training reduces annotation needs. We propose SurgFM-Control, a procedure-aware framework that combines timestamp-aligned frozen-feature caching, a common temporal learner, and adaptation to native procedure labels. This design evaluates representation transfer jointly across surgical procedures, label budgets, and visual perturbations while reusing cached features for downstream experiments. Across three datasets, six frozen encoders reveal complementary strengths: VideoMamba achieves 74.4% and 52.9% macro-F1 on Cholec80 and Microanastomosis, respectively, while DINOv2 reaches 68.8% on AutoLaparo. With one labeled target video, source-initialized temporal learning improves macro-F1 by 5.7% points on average over target only training. DINOv2 also retains 88.7% of clean macro-F1 across the evaluated pixel corruptions, while training sample thinning preserves performance more effectively for long phases than for rapid action transitions. These findings establish a practical basis for matching frozen representations and adaptation strategies to a procedure's annotation budget and temporal structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.