acceptodds
Under review as a conference paper at ICLR 2027

Shared Representations, Native Actions: Learning Generalizable Surgical Motion from Heterogeneous Kinematics

Abstract

Learning surgical tool-motion representations supports action-centric video understanding and autonomous surgery. Existing surgical video foundation models emphasize appearance and semantics, while video-only latent-action models may capture distractors without physical supervision. We present (Surgical Action Representation through Inverse Dynamics Embedding), a surgical video foundation model pretrained through kinematics-supervised inverse dynamics on heterogeneous robot data. Embodiment-specific heads provide kinematic supervision for a shared visual encoder in each robot’s native action space. An auxiliary tool-centric forward dynamics model connects latent actions to visible instrument motion and enables learning from videos without kinematic annotations. Across OpenH, SAR-RARP50, GraSP, and SurgPose, our frozen encoder outperforms the evaluated visual representation baselines on action regression, action recognition, and point tracking. Few-shot adaptation further transfers action prediction to an unseen robot embodiment. In real-robot experiments, augmenting policy training with model-inferred action labels improves needle-pickup success from 60% to 80% and inter-instrument transfer success from 0% to 30% over the limited-demonstration baseline. These results demonstrate the value of coupled forward–inverse dynamics pretraining for transferable surgical representations and the effectiveness of inferred action labels for downstream control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.