acceptodds
Under review as a conference paper at ICLR 2027

Room to Remember: Knowledge-Graph Memory for Long-Horizon Embodied Question Answering

Abstract

An assistant asked about a room it has already left cannot answer from its current image. We study how a vision-language model should read a retained record of those observations. Our benchmark in simulated ProcTHOR houses pairs recorded trajectories with questions about object locations, room visits, temporal order, and accumulated observations. Predicted scene graphs record objects, rooms, relations, and sighting times; we train answerers to read them as text, learned graph features, or both. Training and evaluation use blank answer images to separate memory access from answer-time vision. On 800 questions from 200 test episodes, the textual graph-memory representation improves MomaGraph-R1 accuracy by 6.62 percentage points over a no-memory control across two training seeds, an effect concentrated in route-history questions (+18.9 points) with little change on the rest (). Learned graph features alone show no detectable improvement over no memory on either backbone at either seed. These comparisons favor text within the tested architectures and training schedules. Blind probes fall below chance on the families carrying the gain and well above it on the rest, so aggregate accuracy reflects both retained evidence and question priors. The results support training a text reader as a baseline for episodic memory.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.