acceptodds
Under review as a conference paper at ICLR 2027

HiREx: Hierarchical Retrieval and Active Observation over Visual Memory for Embodied Exploration and Reasoning

Abstract

Embodied exploration and reasoning is the capability to explore an environment to answer questions or reach specified goals. Recent zero-shot methods build a scene representation from observations, over which a vision-language model (VLM) reasons to answer and to decide where to explore. However, existing representations fall short. 3D scene graphs abstract the scene into objects and relations, discarding visual evidence before the task is given. Image-set methods preserve the evidence but retrieve it through detected classes, inheriting their errors. Both leave exploration separate from retrieved evidence. To address these issues, we propose HiREx, a scene memory that organizes the image set under room, view, and object layers, on which both retrieval and exploration operate. Views are selected to preserve object co-visibility, rooms index the views as semantic anchors, and object crops supply fine-grained details. Given a task, retrieval narrows the evidence through the layers, and the level at which evidence is found activates exploration options at that level. Across question-answering and goal-navigation benchmarks, our method outperforms all baselines and shows that memory, retrieval, and exploration each contribute to embodied exploration and reasoning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.