Finale: Reconstructing Local Supervision with Terminal-Anchored Replay for End-to-End Web Agents
Abstract
End-to-end Web development requires implementing and integrating interdependent features. Complete teacher trajectories demonstrate both local implementation and the overall development process. We investigate whether more explicit local supervision from existing teacher demonstrations can improve end-to-end performance beyond complete-trajectory training alone. Identifying these local tasks is difficult because repeated file revisions leave multiple candidate endpoints, while local objectives and action boundaries remain unlabeled. We introduce \method, which reconstructs these tasks using retained file states and successful builds as observable endpoint anchors. It comprises two reconstruction workflows: ArtifactReplay groups file histories contributing to shared functionality retained in the delivered project, while BuildReplay extracts edit phases ending in successful builds. For each unit, \method recovers the repository prestate, synthesizes a local query from recorded target evidence, and uses endpoint-supported teacher actions as supervision for the model. We mix these units with complete trajectories for training. Across the Web and Game sets of ArtifactsBench and Cookie-Bench, \method improves Qwen3.5-35B-A3B-Base to an average score of 47.29, surpassing its official post-trained counterpart by 17.2%. Applied to Qwen3.6-27B, it raises the score from 40.73 to 62.22, exceeding DeepSeek-V4-Pro (59.05). In controlled comparisons, \method outperforms complete-trajectory training by 4.8 points under the same two-epoch schedule and by 17.2 points at comparable processed-token budgets, while surpassing intermediate replay and same-target reweighting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.