acceptodds
Under review as a conference paper at ICLR 2027

Coarse Intent, Fine Physics: Equipping VLM Agents with Latent World Model Foresight for Open-Loop Manipulation

Abstract

Frontier agents such as Claude Code with Claude Opus 5.5 and Codex with GPT-6 Astra are increasingly used out of the box in robotics, yet they cannot reliably anticipate the physical consequences of low-level actions. JEPA-style latent world models predict these consequences quickly, but planners built on them need goal images and exploit optimistic prediction errors, a failure we call hallucination. We give Claude Code access to DDP-WM through two tools, one previewing a proposed plan as keypoint clouds predicted under one hundred perturbations and one running a local cross-entropy search seeded by that plan, and the agent learns to use both from a single exploratory episode. On a Push-T protocol that requires moving the block 61 pixels and rotating it 77 degrees on average within one open-loop plan 2.4 times shorter than human demonstrations, MPC with DDP-WM solves 2 of 21 episodes and the agent alone 7, while the two tools raise this to 14 and 16. Replacing the clouds by single predictions lowers these numbers to 12 and 13, showing that the agent uses the predicted distribution to avoid hallucinated plans.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.