acceptodds
Under review as a conference paper at ICLR 2027

ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

Abstract

Coding agents are usually evaluated on tasks whose desired behavior is stated in an issue, instruction, or test. In web development, however, the specification often exists only as working software, such as a prototype or a comparable product. We study reference-guided software engineering, in which an agent discovers missing behavior by interacting with a live reference whose source is hidden and implements it in an incomplete codebase. We introduce ProgramDistill, a benchmark in which a recorded browser workflow of a running application serves as both a task and its verifier. Its automated task generation pipeline, mine–craft–patch, deletes the code behind each workflow and replays the workflow to check the repair, so no issues or human-written tests are needed. Masking several dependent behaviors controls restoration depth, yielding 4,063 tasks from 26 applications. Across nine frontier agents, performance drops sharply with depth. GPT-6 Astra falls from 100% at depth 1 to 64.0% at depth 8, and every other agent retains less than half of its depth-1 performance. When rebuilding entire applications, the best agent passes only 49.2% of end-to-end workflows. Most unrecovered behaviors (59.2%) were never observed in the reference, suggesting that exploration and validation, not only code generation, are key bottlenecks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.