VTEM: View-Transition Egocentric Memory for Long-Horizon Embodied Reasoning
Abstract
Long-horizon embodied reasoning is fundamental to the reliable deployment of embodied agents, enabling informed decision-making from observations accumulated during long-term interactions. However, existing memory-based approaches often fail to establish effective spatiotemporal associations across observations, causing redundant cross-view content to accumulate in memory and retrieval to return context-misaligned evidence. To address these limitations, we propose the View-Transition Egocentric Memory framework (VTEM), which models view transitions to support traceable memory construction and reliable evidence retrieval. VTEM comprises two core components: the View-Transition Memory Chain (TMC) and Transition-Guided Hierarchical Retrieval (THR). TMC derives spatiotemporal associations from sampled observations and uses them to assess memory value, retaining observations that contribute novel information in a hierarchical scene-view memory chain, thereby enabling efficient memory construction. THR decomposes user questions into hierarchical retrieval targets and association cues, then combines scene-view matching with transition-guided multi-hop reasoning to retrieve fine-grained, context-aligned evidence. Extensive experiments on OpenEQA, ECBench-MultiScene, and EnvQA demonstrate improvements of 7.89, 3.44, and 4.57 points over the strongest baselines, respectively, validating the effectiveness of VTEM for efficient and reliable long-horizon embodied reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.