acceptodds
Under review as a conference paper at ICLR 2027

Repair or Rebuild? What Validators Miss When Agents Revise Executable World Models

Abstract

When the rules of its world change, an agent can repair its executable world model or rebuild it, and a validator decides which revision its planner will trust. To test whether that trust is earned, we executed the confident plans of accepted programs in a pre-registered Crafter-OO audit. Acceptance did not protect them: they failed on new tasks about as often as rejected candidates' plans (19.0% against 23.3%, exploratory), and a pre-registered replication gives 17.3% (10.8–25.0%). Repair solved the most tasks, and collateral regressions, errors at rules the change left valid, came almost only from rebuilds; 136 of the 265 failing plans were shortcuts through a world that exists only in the model (post hoc). The pre-registered strictest gate, perfection on sealed held-out transitions, removed every failure and admitted 7 of the 30 accepted programs. We then turned to code, where no engine defines correct behavior. For 11 of 62 resolved SWE-bench Verified patches, the issue or the unpatched code certifies a counterexample, against 1 of 13 coder-cleared controls. On one selected issue, all 8 top systems ship the same crash that every released test misses, and we prove that agreement among revisions cannot flag a defect most revisions share. A detector anchored on the unpatched code raised fewer false flags than a matched auditor on every backbone (pre-registered). Gate revisions on a regression check and a sealed check, never on agreement alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.