acceptodds
Under review as a conference paper at ICLR 2027

Closing the Act-Predict-Evaluate Loop with a Shared Video-Pretrained Representation for Robotic Manipulation

Abstract

Video-pretrained representations provide rich visual and temporal structure for embodied learning, but acting, prediction, and evaluation are still typically developed in separate representation spaces. This separation limits model-based policy improvement because predicted future states cannot be directly consumed by the policy and value model. We introduce Pivot, a framework that instead uses a frozen video-pretrained representation as the shared state space of the entire Act-Predict-Evaluate loop. The policy reads this state, action-conditioned dynamics predicts its successor, and a value model evaluates the predicted state under the instruction, enabling recursive latent imagination without observation reconstruction or representation translation. To make this generic video representation effective for manipulation, we introduce task-aware supervision at two levels. First, we train task-relevant latent dynamics using a semantic affordance prior derived from image-instruction attributions, emphasizing the interaction-critical changes. Second, we learn a task-regularized value function with temporal-difference learning and dense image-instruction supervision, providing progress estimates for imagined states. These estimates allow latent rollouts to directly guide policy improvement. On CALVIN, Pivot improves the imitation policy's average successful sequence length from to . On a five-stage real-robot task, it increases the average number of completed stages from to , with controlled ablations showing that both forms of task-aware supervision contribute to the improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.