Structured Distillation of Web Agent Capabilities Enables Generalization
Abstract
Web agents powered by frontier models are increasingly able to navigate complex websites, but they rely heavily on proprietary APIs, which increases costs and hinders for local deployments. To solve this, smaller models can be hosted locally, but are often less capable. We take a step towards bridging this gap by introducing AGENT-AS-ANNOTATORS, an automatic framework for distilling the capabilities of frontier models to navigate websites into smaller models, which substantially improves their abilities to solve tasks on websites. The framework is designed to replicate how humans would collect trajectories for training web agents; but the entire pipeline is automated using modular LLM components based on frontier models. Using Gemini 3 Pro as the teacher model, we generate 3000 trajectories across six web environments from WebArena, which we filter and use to train a student (Qwen-3.5-9B) with only supervised finetuning (SFT). We find that the resulting agent can transfer beyond the environments used in its training. Notably, on WorkArena L1, success rises from 33.3% to 51.5%, and performance also improves across three additional benchmarks, each testing a different type of generalization. On WebArena, it achieves a success rate of 41.5% despite never seeing any of the benchmark tasks, which nearly doubles the best same-size agent (Go-Browse, 21.7%). Through ablations, we find that the model benefits from each of its components (personas, judge, hints) and reasoning traces. Our results show that structured trajectory synthesis from a single teacher can produce competitive yet locally deployable web agents. We release the code, data and family of models (ranging from 2 to 9B parameters) publicly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.