CARE-VLA: Context-Aware Retrieval of Execution Experience for Agentic VLA Planning
Abstract
Agentic vision–language–action (VLA) systems use external planners and execution memories to support frozen control policies. However, memories organized around complete tasks or episodes do not directly align with local planning decisions, and irrelevant or redundant experience can distract the planner. We propose CARE-VLA, a training-free framework for context-aware retrieval of step-level execution experience. CARE-VLA constructs a read-only multimodal experience bank from reference trajectories, linking each primitive execution to its preceding context, outcome, and before–after observations. At each decision boundary, it filters records by execution-context compatibility, fuses task-and-context and visual rankings, and selects a small set of relevant records with diverse next primitives and distinct source episodes. The selected experience forms a Memory Card that supplements the planner's original context without updating planner or VLA parameters. On LIBERO-Pro and RoboTwin C2R, CARE-VLA achieves success rates of 80.3% and 61.6%, respectively, exceeding the published Harness VLA results with the corresponding Codex planners by 8.2 and 3.6 percentage points. Compared with Harness VLA, CARE-VLA reduces the average number of primitive executions by 8.9% on episodes successfully completed by both methods. These results support step-level experience retrieval as a means of improving frozen agentic VLA planning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.