acceptodds
Under review as a conference paper at ICLR 2027

BoundaryLoop: From Failure Evidence to Persistent Web Capabilities

Abstract

An episodic WebAgent repair becomes a reusable capability when its executable behavior survives new task parameters and later runs. We study how failure evidence supports that transition. BoundaryLoop turns failed WebAgent episodes into browser-tested instance workflows, compares repairs to extract parameterized family procedures, and admits those procedures to executable memory after transfer and boundary validation. The WebAgent supplies task, action, DOM, and browser-state evidence; an auxiliary CodeAgent synthesizes and tests the correction. In a frozen WebArena Shopping Admin comparison with 84 runs per condition, the Base WebAgent records 2/84 execution successes and 0/84 clean completions. Static repair passes the evaluator in 59/84 runs but has parameter pollution in 84/84, whereas promoted family workflows complete 84/84 runs without recorded pollution or drift. Compact aligned evidence matches the 6/24 accepted-repair yield of full raw evidence with 12.7% fewer input tokens. A frozen family procedure transfers to 3/3 original instances excluded from abstraction and 5/5 post-freeze generated instances. In a separate later-runtime study, stored workflows pass 33/33 runs versus 0/33 for the matched BaseAgent. These results locate capability acquisition in cross-instance invariance and validated reuse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.