ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks
Abstract
Coding agents are usually evaluated on tasks whose desired behavior is stated in an issue, instruction, or test. In web development, however, the specification often exists only as working software, such as a prototype or a comparable product. We study reference-guided software engineering, in which an agent discovers missing behavior by interacting with a live reference whose source is hidden and implements it in an incomplete codebase. We introduce ProgramDistill, a benchmark in which a recorded browser workflow of a running application serves as both a task and its verifier. Its automated task generation pipeline, mine–craft–patch, deletes the code behind each workflow and replays the workflow to check the repair, so no issues or human-written tests are needed. Masking several dependent behaviors controls restoration depth, yielding 4,063 tasks from 26 applications. Across nine frontier agents, performance drops sharply with depth. GPT-6 Astra falls from 100% at depth 1 to 64.0% at depth 8, and every other agent retains less than half of its depth-1 performance. When rebuilding entire applications, the best agent passes only 49.2% of end-to-end workflows. Most unrecovered behaviors (59.2%) were never observed in the reference, suggesting that exploration and validation, not only code generation, are key bottlenecks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.