acceptodds
Under review as a conference paper at ICLR 2027

When to Perceive, Remember, Imagine, or Act: Autonomous Decision for Embodied Agents

Abstract

Embodied agents must decide what to do from incomplete and potentially conflicting visual evidence. Existing approaches offer complementary mechanisms for visual reasoning, memory retrieval, and predicting action outcomes. Some recent methods also adaptively choose when to gather imagined evidence, but coordinating perception, memory, imagination, and execution remains a broader decision problem. Without such coordination, an agent may overlook relevant details, ignore past experience, or rely on imagined outcomes that are inconsistent with reality. To address this problem, we propose AutoPRIA, a framework that Autonomously organizes embodied decision-making around four fundamental operations: Perceive finer details of the present, Remember the past, Imagine the future outcome of a candidate action, and Act to obtain a real observation from the environment. To the best of our knowledge, AutoPRIA is the first framework to bring all four operations together in one embodied agent, so that the agent decides what to know together with what to do. Under this framework, we build an agent with tool-augmented perception, dual memory, action-conditioned imagination with a fixed world model, and real actions whose observations refresh the information state and check earlier predictions. The policy is trained with supervised initialization followed by reinforcement refinement. Across 4 embodied tasks (active recognition, image-goal navigation, active embodied question answering, and robotic manipulation), AutoPRIA outperforms the Qwen3.5-9B baseline with the Wan2.2 world model on every task, with an average relative improvement of 17.7%, while executing fewer actions on AR, IGNav, and RM.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.