Learning to Recover: Iteratively Internalizing Agent Corrections into VLAs for Robotic Manipulation
Abstract
Vision-language-action (VLA) models acquire manipulation skills through large-scale training, while reasoning agents can use tools and policies to handle situations beyond those skills. Yet successful assistance does not itself improve the underlying policy, leaving subsequent executions dependent on renewed online reasoning. We present Learning to Recover (L2R), an agent-to-policy learning framework that iteratively internalizes agent corrections into reusable recovery capabilities. At natural failures, the agent constructs and verifies corrections; lightweight offline updates then teach the policy to reuse them without online agent assistance. Rather than imitating entire intervention records, L2R selectively learns valid corrective actions and when to invoke them through conditional low-rank adaptation of a frozen VLA and a lightweight local controller. Further rounds target failures of the updated policy, allowing recovery learning to follow the policy's changing behavior. On ten VLABench task families across five tracks, a single round trained only on in-distribution corrections achieves 52.0% conditional recovery, compared with 31.8% for standard LoRA. Three rounds improve policy-only task success from 12.6% to 27.4% on ten RoboDojo tasks and from 33.3% to 68.3% on three real-world tasks. These results show how limited agent corrections can be progressively internalized into autonomous recovery capabilities. Videos are available on our project page.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.