acceptodds
Under review as a conference paper at ICLR 2027

SurgJEPA: Structured Context for Multi-Horizon SurgicalWorkflow Anticipation

Abstract

Anticipating surgical workflow requires integrating visual observations with procedural context while accounting for the substantial nuisance variation present in future frames. Rather than predicting future pixels directly, we model future surgical scenes in a learned latent representation space. We introduce SurgJEPA, a tool-conditioned joint embedding predictive architecture that forecasts distributions over latent surgical states at 5, 30, 60, and 300 seconds. A causal context encoder integrates historical visual features, phase information, and temporally aligned tool observations, together with explicit missingness indicators that enable robust inference when tool data are incomplete. For each prediction horizon, a horizon-conditioned predictor generates four weighted latent hypotheses, while an exponential-moving-average target encoder provides self-supervised representations of future observations. In a five datasets study comprising 301 cases with case disjoint splits, tool-conditioned training improves dataset macro future phase macro-F1 over a matched tool free configuration by 10.2%, 9.4%, 9.1%, and 5.0% at 5, 30, 60, and 300 seconds, respectively, while reducing phase negative log-likelihood at every horizon. Representation level analyses further show that the learned contextual state supports future-phase decoding accuracies of 83.5%, 74.3%, 72.7%, and 53.5% across the four horizons, and that the predicted mixture representation retains complementary information about future workflow state. A controlled pilot additionally demonstrates that latent temperature substantially affects the geometry of the predicted representation, increasing pooled effective rank from 2.80 to 17.54. Together, these findings demonstrate the value of structured, tool conditioning multi-horizon surgical workflow anticipation and clarify how future phase information is distributed across the learned latent representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.