acceptodds
Under review as a conference paper at ICLR 2027

Verified World-Model Repair for Long-Horizon Interactive Agents

Abstract

Long-horizon interactive agents can fail because a false belief about the world survives the very action that should disprove it. We introduce verified world-model repair, a framework that maintains a structured belief state, localizes action-conditioned prediction errors, converts them into scoped repair programs, and commits a repair to memory only after independent validation. A stepwise recovery specialist learns complete successful recovery trajectories rather than isolated next-action corrections. We also construct the World Evolution Dataset (WED), a family-disjoint dataset and paired-recovery protocol derived from 10,000 interaction episodes and 162,756 transitions. Across five random seeds on ALFWorld failure checkpoints, the full system recovers 169 of 185 cases (91.35%), compared with 109 of 185 (58.92%) for a strong goal-state controller. Learning only the first corrective action reaches 44.14%, whereas learning the full recovery trajectory reaches 90.99%. The method remains effective on held-out scenes, objects, and task templates, and reaches 93.69% on Qwen3-8B after development-only adaptation, while limiting harmful corrections to 3.6%. On frozen WebArena and OSWorld subsets, it improves programmatic task success over the strongest reflection baseline by 6.10 and 6.00 percentage points, respectively. These results indicate that reliable long-horizon recovery requires verifying a correction and carrying it through the remaining task, rather than merely revising the next action.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.