acceptodds
Under review as a conference paper at ICLR 2027

ANTE: Anticipate, Then Execute - Learning Internal World Models for Code Agents from Verified Beliefs

Abstract

Rapid advances in large language models have made software engineering agents capable of tackling complex coding tasks. Within a containerized environment, a code agent receives tool responses or execution results, and determines its later actions by those earlier results. However, most existing code agents never look ahead: their policies do not anticipate the consequences of actions, and the terminal reward barely attributes a failure to certain earlier steps, so an ill-judged edit can corrupt the repository state irreversibly before any signal arrives. Recent advances in code world models seek to tackle these challenges, yet they are learned offline or frozen before certain training stages, reduced to an auxiliary loss on observation tokens. We present ANTE (Anticipate, Then Execute), an agentic RL framework that verifies the policy's pre-action predictions by real execution and turns them into per-step credit for its actions and trust-or-abstain decisions. With an internal belief stated before each action, Belief-Anchored Policy Optimization (BAPO) turns the settled predictions into per-step advantages by anchoring credit on the policy's own forecasts. To turn the trained forecasts into better actions at inference, thereby facilitating effective test-time scaling, ANTE-TTO selects among candidate actions without an external scorer. ANTE raises Qwen3-32B from 25.0% to 57.2% pass@1 on SWE-bench Verified, 6.2% above a matched GRPO baseline trained without the belief state; its world model's accuracy climbs from 44% to 69% during RL, and the learned foresight capability transfers across harnesses and out-of-distribution domains. In summary, we aim to build a code world model that is omniscient over both time and environment state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.