FrontierGuard: Learning Executability Frontiers for Structured Web and API Agents
Abstract
A checkout agent can face the same visible submit control twice: before shipping fields are complete the action is illegal, while one successful form event later it is executable. API agents encounter the analogous transition when an earlier response produces a required identifier. FrontierGuard learns this moving boundary as an executable-state quotient over grounded interaction events. Histories are merged only when they preserve action-feasibility profiles and event-conditioned successors; the resulting finite-state controller maps each proposal to allow, block, or one read-only query before any side effect. Across 1,152 WebArena and ToolBench-API tasks, FrontierGuard reduces GPT-4o WebArena illegal actions from 23.4% to 10.7% (54.3% relative), raises success from 18.7% to 24.5%, and cuts dead ends from 4.2% to 1.9%; on ToolBench-API it adds 3.5 success points beyond schema-constrained decoding. At the same 8.4% query rate, the quotient adds 1.3 WebArena success points and removes 1.5 illegal-action points relative to a same-predicate GRU. Pooled leave-one-family-out evaluation retains the strongest listed operating point on both benchmarks, showing that a learned execution boundary can convert local interface evidence into more completed trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.