Runtime Recovery for Frozen Generative Robot Policies under Long-Tailed Distributions
Abstract
Generative behavior cloning can produce diverse action distributions and represent complex, multimodal robot behaviors. However, robot execution is also governed by physical dynamics, including contact, friction, and object motion. Small action errors can be amplified during execution, leading to substantial deviations in the robot state, especially when the robot enters long-tail states that are weakly covered by the training data. A common approach is to sample multiple candidate actions from the original policy and select a more suitable one for execution. However, if the policy itself does not cover successful actions at the current state, a limited increase in the number of samples may still fail to identify an effective action. To address this issue, we propose TailGuard, a runtime recovery method for generative policies in long-tail states. TailGuard keeps the original policy frozen and uses previously successful rollouts to characterize task-relevant successful regions. When execution deviates from these regions, an external controller, together with a learned short-horizon dynamics model, corrects the robot behavior and guides it back toward states from which successful execution is feasible. Control is then handed back to the original policy once it can reliably continue the task while maintaining action continuity. Across four robotic manipulation benchmarks, TailGuard requires neither additional demonstrations nor failure data, improving the average closed-loop success rate from 49.1% to 70.4% and outperforming retrieval- and residual-based recovery as well as recent recovery methods for diffusion policies.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.