Shared Representations, Native Actions: Learning Generalizable Surgical Motion from Heterogeneous Kinematics
Abstract
Learning surgical tool-motion representations supports action-centric video understanding and autonomous surgery. Existing surgical video foundation models emphasize appearance and semantics, while video-only latent-action models may capture distractors without physical supervision. We present (Surgical Action Representation through Inverse Dynamics Embedding), a surgical video foundation model pretrained through kinematics-supervised inverse dynamics on heterogeneous robot data. Embodiment-specific heads provide kinematic supervision for a shared visual encoder in each robot’s native action space. An auxiliary tool-centric forward dynamics model connects latent actions to visible instrument motion and enables learning from videos without kinematic annotations. Across OpenH, SAR-RARP50, GraSP, and SurgPose, our frozen encoder outperforms the evaluated visual representation baselines on action regression, action recognition, and point tracking. Few-shot adaptation further transfers action prediction to an unseen robot embodiment. In real-robot experiments, augmenting policy training with model-inferred action labels improves needle-pickup success from 60% to 80% and inter-instrument transfer success from 0% to 30% over the limited-demonstration baseline. These results demonstrate the value of coupled forward–inverse dynamics pretraining for transferable surgical representations and the effectiveness of inferred action labels for downstream control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.