acceptodds
Under review as a conference paper at ICLR 2027

ReAlign-VLA: Deviation-Guided Recovery for Off-the-Shelf Vision-Language-Action Models

Abstract

Vision-Language-Action models have shown strong potential for language-conditioned robotic manipulation, but reliable deployment requires the ability to recover from execution failures, not merely to generate actions. In practical manipulation, routine visual uncertainty, contact errors, and action drift can turn small deviations into task failure, while recovery strategies based on VLA fine-tuning or auxiliary models limit scalability in open-ended and changing environments. We propose ReAlign-VLA, a model-update-free recovery framework that couples internal-state runtime monitoring with deviation-conditioned intervention. Rather than training a separate failure detector or invoking an external VLM planner, ReAlign-VLA constructs a progress-conditioned profile of successful executions from the frozen VLA's hidden states. During deployment, it extracts hidden states through a lightweight forward hook and computes a sliding-window normalized deviation score to identify when the current execution departs from successful behavior. The deviation score then gates a tiered recovery mechanism, escalating from action-level correction for physical execution errors to deterministic sub-goal refinement for semantic or planning ambiguity. Progress-aware locking and early abort make the intervention conservative, preserving near-successful trajectories while discarding attempts that remain far from the success profile. Experiments on four LIBERO benchmark suites show that ReAlign-VLA consistently improves an unmodified OpenVLA baseline in both task success and execution efficiency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.