REVIVE: Inference-Time Return with Phase-Aware Controller for VLA Models
Abstract
Vision-Language-Action (VLA) models have shown strong general-purpose robot control by leveraging large vision-language models. However, imitation-learned VLA policies remain brittle in multi-step manipulation: small execution errors can compound off the demonstrated trajectory, and the policy may silently advance to a later phase without completing the current subgoal, which we formalize as . Existing recovery approaches often rely on additional human interventions, reward supervision, or offline retraining. We propose REVIVE, an execution-time recovery framework that improves VLA success without retraining the underlying policy. Instead of learning explicit recovery from every failure, REVIVE intervenes before local deviations escalate. An external Return Controller monitors phase-local progress, boundary evidence, and action-effect signals to detect phase-level stalls; upon detection, it performs phase-local rollback to the latest valid phase anchor, clears the current action chunk, and triggers replanning. Failure-directed decoder adaptation generates alternative actions, while failure memory ranks candidates to avoid repeating past mistakes. Experiments on simulated manipulation benchmarks and two real-robot long-horizon tasks show that REVIVE improves task success, including under out-of-distribution task configurations and object changes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.