acceptodds
Under review as a conference paper at ICLR 2027

Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning

Abstract

Synthetic web environments can scale agent training, but plausible pages may hide unreachable goals, contradictory records, or invalid state updates. These defects confound policy failure with environment failure. We study environment reliability as a source of supervision for compact web agents. Our framework represents a website as a scaffold of pages, navigation, backend records, tasks, and state-change markers. Bounded verification and local repair establish executable task traces; marker-constrained updates and backend-derived rewards connect visible actions to persistent state. Across 500 environments in six domains, task feasibility increases from 48.6% to 94.8%. With terminal rewards fixed, verification raises held-out PPO success from 19.4% to 42.6%; adding dense state rewards reaches 58.7%. Controls separate these effects from feasible-task filtering and reward granularity. The resulting policies also improve on WebArena-compatible tasks, WebShop, and MiniWoB++ through a shared DOM interface, without policy-side LLM calls. The evidence supports bounded, DOM-grounded interaction learning while exposing remaining limitations in noisy web dynamics and verification coverage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.