TrajSmith: Reusable Web Tools from Trajectories through Exploration and Verification
Abstract
Web agents repeatedly plan browser procedures that earlier trajectories already contain. This repetition increases inference cost and exposes smaller models to errors over long action sequences. Converting trajectories into reusable tools can reduce these demands, but a working source trajectory does not establish when the resulting tool will apply on a changing website. We present TrajSmith, a trajectory-based method for synthesizing web tools with explicit reuse conditions. A strong model actively explores the structure around a successful or failed trajectory, inspects related controls and input options, and compiles the observations into atomic and composite tools. An applicability stage challenges candidate tools with changed arguments, boundary cases, neighboring intents, and empty-result queries. Structural checks and an expected result type screen the returned evidence, and a Tool-Use Contract records argument domains, use conditions, and rejection rules. At deployment, a weak model reads the contract to select a tool and fill its arguments; contract checks determine whether the proposal proceeds to execution or fallback. On a 159-task Online-Mind2Web subset selected for tool coverage, the better of two paired runs raises Qwen3-4B success from 18.2% to 30.8% while using fewer recorded agent-loop steps. On WebArena's 763 single-domain tasks, success increases from 19.8% to 24.5%. Additional shared-budget comparisons show gains on source-overlapping tasks for GPT-4o-mini and Qwen3-VL-4B. These results support reuse of recorded procedures across repeated runs and several weak-model configurations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.