ERAct: Towards Reasoning over Memorization via Semantic Experience Rationale Informed Planning
Abstract
Vision-Language Models (VLMs) have shown strong potential for embodied planning, yet their performance can degrade substantially under distribution shifts. While accumulated experience provides a scalable source of external knowledge, raw trajectory-based experience provides concrete behavioral references but leaves the decision logic underlying actions implicit, making it difficult to reuse flexibly when current executions diverge from past trajectories. Based on this insight, we propose Experience Rationale Informed Planning (ERAct), a rationale-driven embodied planning framework that transforms past experiences into explicit reasoning knowledge and transfers decision-relevant reasoning cues to novel situations. ERAct consists of a Semantic Experience Rationale Retriever (SERR), which performs coarse-grained retrieval of reasoning-relevant experience instances, and a Rationale-Informed Fusion Network (RIFN), which conducts fine-grained dual-path relevance modeling to extract useful reasoning cues and adaptively inject them through a confidence-aware residual mechanism. Extensive experiments on the EB-ALFRED benchmark within EmbodiedBench demonstrate that ERAct improves both instruction understanding and environmental reasoning, consistently outperforming prior agents. Further analyses show that ERAct exhibits stronger reasoning generalization and more effective experience utilization than direct trajectory reuse.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.