MUSHROOM: Benchmarking Long-Term Spatial Reasoning in Multi-Room Environments
Abstract
Long-term spatial reasoning requires vision-language models (VLMs) to consolidate partial observations into spatial knowledge and update it as spatial states change. We introduce MUSHROOM, a benchmark for long-term spatial reasoning in multi-room environments, comprising 230 egocentric videos that traverse multiple rooms, where room boundaries disrupt spatial continuity. We design tasks along two complementary axes: Static Spatial Reasoning within a stationary configuration, and Dynamic Spatial Reasoning over state changes. We further propose SPORE, a training-free pipeline that organizes a single video into a structured representation: an object-centric spatial graph. Our evaluation demonstrates that contemporary VLMs have fundamental limitations in cross-region reasoning and updating spatial states, while SPORE substantially improves both static and dynamic reasoning, with consistent gains on real-world indoor benchmarks. Further analysis shows that the spatial graph links temporally distant observations and stabilizes geometric estimation, yet cross-region reasoning falters under frequent transitions, a pattern reminiscent of the doorway effect.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.