Trajectory-aware Experience Retrieval with Contrasting Success and Failure Memories for Embodied VLM Agents
Abstract
Vision-language models (VLMs) have improved embodied agents, but errors in perception, planning, and execution often accumulate in unseen environments. Existing memory-based methods do not fully address this limitation: they retrieve experiences using the task instruction and current observation, favoring goal-level similarity over the agent’s actual execution history. Yet the guidance an agent needs depends on that history, as continuing a successful trajectory requires different guidance from recovering after an error. To address these limitations, we propose Trajectory-aware Experience Retrieval for embodied Agents (TERA), a training-free framework that retrieves guidance based on the trajectory executed so far. TERA summarizes each multimodal trajectory, infers failure causes directly from failed executions, and distills situation-conditioned guidance into separate success and failure memory banks. At inference time, TERA queries both banks using a summary of the current trajectory, enabling the planner to contrast guidance on how to proceed with guidance on what to avoid and thereby improve action selection. Across 16 VLM backbones, TERA improves the mean success rate on EB-Manipulation by up to 7.9% and improves Skill Match, Entity Match, and Exact Match on VLABench by up to 13.2%, 12.9%, and 9.4%, respectively. Ablation studies show that both trajectory-aware retrieval and the separation of success- and failure-derived guidance contribute to these gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.