JANUS: Spatial Belief and Evidence Imagination for Embodied Active Perception
Abstract
Active perception is an essential capability for embodied agents, providing the ability to selectively acquire the visual evidence required for downstream decision-making. Recent approaches leverage multimodal large language models (MLLMs) to decide where to observe next. However, without persistent memory and explicit spatial representations, these approaches struggle to maintain spatial consistency across observations and identify informative regions for further exploration. In this work, we seek to unlock the memory and associative reasoning capabilities of MLLMs, investigating whether an agent can efficiently accomplish visual perception tasks through a sequence of selective egocentric observations. To this end, we introduce JANUS, a framework that couples persistent spatial belief with imagination-guided next-view decisions. Specifically, we develop a Hierarchical Spatial Belief module that organizes accumulated real observations into a hierarchical representation over scenes, regions, objects, and their attributes, while preserving confidence estimates and links to the supporting visual evidence. We further introduce a Belief-Conditioned Evidence Imagination module that conditions on the current spatial belief to anticipate informative candidate observations, using predicted latent evidence to reactivate relevant past observations and guide next-view selection. Together, these modules connect accumulated observations with future information needs, enabling belief-guided decisions about where to observe next. Experiments on challenging perception benchmarks demonstrate the effectiveness of JANUS for embodied active perception.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.