The Agent Is in the Details: What Matters in Web Agent Harness Design?
Abstract
A web agent harness determines what the model sees, what it retains across steps, and how model outputs become browser actions. Harness designs are often introduced inside complete systems with their own models and setups, so reported gains do not show which design to adopt. We build a configurable framework on GenericAgent from AgentLab, implement 30 configurations in six functional categories, and compare each with its setting-specific baseline on WebArena-Lite with Qwen3.5-9B and OpenApps with Qwen3.5-27B. The two highest-scoring configurations in the reported catalog results are multi-turn histories in both settings: full observation history keeps earlier page bodies, whereas observation masking replaces them with placeholders; both retain past model responses. Full history and masking reach success rates of 26.06% and 23.64% against a 15.76% baseline on WebArena-Lite, and 57% and 31% against 19% on OpenApps. On WebArena-Lite, masking trails full history by 2.42 percentage points and uses about a third fewer action-model tokens than the baseline; on OpenApps, it trails by 26 points. Only one tested combination outscores the constituent configurations alone: full history with execution-error feedback on OpenApps. We also evaluate one fixed workflow library, distilled from successful trajectories on the evaluated task set, across five WebArena-Lite model settings. Ordered by baseline success, the settings show declining workflow increments, from +7.88 points with Qwen3.5-4B to −1.82 with GPT-5.6 Sol. These results favor evaluating multi-turn history first, comparing combinations with the constituent configurations alone and the unchanged baseline when selecting by success rate, and evaluating this library on the target model. We will release the framework code and configurations as an extensible platform for web agent harness research, supporting new designs and reuse of existing comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.