WEBACT: Test-Time Learning of Verifiable Action Interfaces for Web Agents
Abstract
A web agent that has just finished a routine cannot call that routine again. Its next task begins again from browser primitives, and the agent works out the same clicks and keystrokes a second time. Making the routine callable via skill induction solves only part of the problem. A routine that worked once can be invoked on a page where it cannot work, and even on the right page it can run to the end without producing the effect that the task needed. We study test-time learning of verifiable action interfaces, where an agent learns reusable operations together with the conditions under which they should be used. Our framework, Webact, learns an ActSpec for each routine: a parameterized plan, a rule for finding its targets on the current page, a condition for when it may be invoked, and a check on the effect it should produce. Every one of these fields is taken from the trace that the routine was learned from. Successful attempts add operations, failed attempts add constraints, and each invocation outcome adjusts which operations the planner prefers, all without changing model weights. On WebArena tasks, the learned action interface improves task success by 32.8% relative to a primitive-action baseline while using 42.9% fewer tokens per task, and the gain holds on live websites. Removing either of the two learned conditions lowers success. Further trace analysis revealed the consequences of applying a wrong skill: once the planner selects an operation on the wrong page, it adopts a sub-goal that the task never had, and rejecting the call does not remove that sub-goal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.