WebDevHorizon: Evaluating Coding Agents on Complete Web Feature Delivery
Abstract
Driven by rapid advances in large language models, autonomous coding agents are increasingly being applied to full-stack Web development. Existing benchmarks, however, typically evaluate either the resolution of isolated repository-level issues or the greenfield development of Web applications from scratch. In practice, developers regularly implement user-facing features within established Web codebases. Delivering such features is inherently a long-horizon, cross-layer endeavor: an agent must understand and modify frontend interfaces, backend logic, and often database schemas, while keeping states synchronized across these layers through iterative test and runtime feedback. Yet, it remains uncertain how reliably autonomous agents can complete such full-stack feature implementation tasks. In this paper, we introduce WebDevHorizon, a benchmark for long-horizon, full-stack Web feature delivery. It comprises 114 tasks obtained from 52 releases across 24 public repositories. Each task corresponds to a product feature shipped by the repository maintainers, with its behavioral requirements reconstructed from the project's development history. To evaluate feature-level behavioral completeness, WebDevHorizon maps these requirements to executable checks of component behavior, API contracts, persistent state, and browser workflows as needed. The environments and test suites are human-verified, and a task is resolved only when all target and regression checks pass. Evaluating eight frontier LLMs with mini-swe-agent highlights the difficulty of long-horizon delivery: across trajectories averaging 56–161 agent steps per task, models resolve 40.35%–63.16% of tasks, led jointly by Opus-5 and GPT-5.6-Sol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.