Reflect, Recover, Internalize: Self-Improving VLAs from Execution Feedback
Abstract
Vision-language-action (VLA) models are usually trained from fixed demonstrations and discard the consequences and corrections observed during deployment. We introduce RRI-VLA, a reflect-recover-internalize framework that uses execution feedback for both immediate recovery and persistent improvement. At each action boundary, a frozen multimodal reasoner summarizes the preceding chunk as an outcome, running episode memory, and next adjustment. These states condition the next action, while a governed orchestrator combines semantic, physical, and learned evidence to authorize one bounded corrective chunk. An offline verifier then converts completed rollouts into verified failure-recovery records that train an action-conditioned Recovery Critic across deployments. After one-time feedback-interface initialization, only the critic is updated and the action-producing VLA remains frozen. On LIBERO-Plus and LIBERO-PRO, RRI-VLA reaches 92.47% and 57.84% task success. Relative to online-only recovery, internalization adds 3.12 and 5.64 points while retaining 95.30% success on unperturbed tasks; governed recovery resolves 74.2% of eligible failures with 1.10% harmful flips. Ablations identify verified experience, causal feedback alignment, and bounded governance as the key sources of improvement. Execution feedback thus provides a shared interface for within-episode recovery and policy-preserving learning across deployments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.