Correct the Action, Not the Policy: Online Recovery under Dynamics Shifts
Abstract
Changes in object dynamics can cause a pretrained manipulation policy to fail even when object geometry and task goals remain unchanged. We study whether such failures can be recovered during execution without updating the policy itself. We introduce a model-based action correction approach that estimates effective dynamics from interaction records, trains a lightweight predictive model under the estimated dynamics, and modifies only a low-dimensional subset of the frozen policy’s actions. During execution, candidate corrections are evaluated from the current observation and selected to keep predicted motion close to a successful trajectory produced by the same policy under reference dynamics. Across shifts in object mass and contact friction, the method raises success from 0–4% to 52–78% in rope routing and T-pushing, and from 65% to 78% in toy packing. A separate physical rope experiment further shows that a simulation-selected correction can recover an execution that fails after a change in object properties. These results show that some failures caused by dynamics shifts can be recovered through execution-time action correction rather than policy retraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.