acceptodds
Under review as a conference paper at ICLR 2027

Learning To Plan with Agentic Imagination

Abstract

A single push in Sokoban can make a puzzle unsolvable. Avoiding such errors requires anticipating consequences before acting, which is especially hard when feedback is sparse and actions cannot be undone. We introduce Agima, a training framework that teaches an agent to imagine before it acts. A single omni model both proposes actions and generates images of where they lead, then treats these imagined frames as observations. It can follow a plan forward, test an alternative move, or compare competing plans, all before touching the real environment. Supervised fine-tuning teaches these forms of imagination, and reinforcement learning then optimizes the model's textual planning decisions using only binary task-success rewards. Across four environments spanning Maze2D, Sokoban, Pacman, and PushT, Agima plans in harder or unseen configurations excluded from training. With supervised fine-tuning alone on two-box Sokoban, it solves 40% of four-box puzzles while seeing the real board only once every five moves; an action-only planner fine-tuned from the same backbone solves 2%. After RL, Agima solves more held-out Sokoban puzzles and adapts how it uses imagination: it drafts a second plan mainly when it foresees the first one failing, and branches more often at critical decisions. An agent that can picture where its actions lead can learn when and where to look ahead, in other words, to plan.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.