Online Robot Policy Refinement via Dynamics-Value Decomposition
Abstract
Pretrained imitation policies can solve complex manipulation tasks, yet small errors can still cause failure in precise manipulation. Online reinforcement learning (RL) can correct such failures, but most policy refinement methods rely on -functions whose offline initialization can be unreliable for the nearby unseen actions introduced during refinement. Rather than directly pretraining a -function, we compose a pretrained latent dynamics model, adapted with task-specific offline transitions, with a state-value model initialized through task-progress supervision. We call this dynamics-value decomposition: an action chunk is evaluated by predicting its latent successor and scoring the predicted state. This composition provides an initial action evaluator that generalizes locally to action perturbations during refinement. During online RL, new transitions update the dynamics model, TD learning calibrates the value model to environment returns, and policy gradients propagate through their composition. Instantiated with DINO-WM and RoboMeter, our action evaluator replaces direct -functions in residual and latent-steering policy refinement methods. Across precise manipulation tasks in simulation and the real world, our method improves online sample efficiency over the corresponding -based methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.