acceptodds
Under review as a conference paper at ICLR 2027

: Advantage based Regularization for Online Policy Repair

Abstract

Vision-Language-Action (VLA) models are popular for zero-shot manipulation across diverse tasks, but their effectiveness is conditioned on exposure during training to **both** the deployment embodiment **and** the specific task domain. We observe that embodiment-grounded VLA policies frequently exhibit _partial competence_ in unseen tasks, even in those closely related to tasks seen in training. These policies complete a task in some non-zero proportion of attempts but frequently make non-recoverable errors, preventing zero-shot deployment of the policy. To address this we propose Advantage-based Regularization for Online Policy Repair (), a method that learns to correct the critical errors that lead to task failure when deploying VLA policies in unseen tasks. Our method requires neither expert demonstrations, reward shaping, nor VLA finetuning. bootstraps a residual action correction solely from the rollouts of a frozen base policy. An advantage regularizer grounds a critic on expected returns from the base policy before finding local corrective actions that would improve upon executing the VLA policy directly. We evaluate the ability of to repair policies on both small scale goal conditioned control environments and VLA policies in unseen manipulation and decision making tasks. Across five held-out LIBERO-90 tasks with 21.8–54.7% base success, improves final success to 55.6–96.3%, outperforming residual RL without advantage regularization on every task. excels at correcting acute critical errors that lead to policy failure, increases the sample efficiency of policy repair, and accommodates multi-step corrections for perturbed configurations of tasks seen during training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.