acceptodds
Under review as a conference paper at ICLR 2027

PolicyGround: Recovering from High-Level Failures by Probing Frozen Hierarchical Vision-Language-Action Models

Abstract

Hierarchical vision-language-action (VLA) systems support long-horizon manipulation by decomposing a task into language subtasks executed by a low-level policy. After a high-level failure, however, a semantically correct recovery instruction may still be unsupported by the particular policy that must execute it. Assessing this support is difficult for multi-step recovery: later policy responses depend on states produced by earlier actions, so policy support cannot generally be assessed from a single policy query, while physically trying each candidate defeats one-shot recovery. We introduce \method, a training-free framework for recovering from high-level failures in frozen, black-box hierarchical VLAs. \method first derives a policy-independent recovery goal from the recorded failed episode. Its key mechanism, Chain-of-Probe, then connects high- and low-level target-policy responses through proxy contexts retrieved from the failed episode, constructing a proxy multi-step response before physical execution. This policy-specific evidence is used to evaluate and revise the recovery instruction toward one supported by the target policy. We evaluate \method with RoboMME-FR, a one-shot physical recovery protocol over ten history-dependent RoboMME tasks. Across these tasks, \method improves micro-averaged recovery success from 7.7% to 23.4%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.