Manufacturing a Progress Signal: Self-Evaluating Coding Agents for Long-Horizon Program Synthesis from Scratch
Abstract
Coding agents are increasingly undertaking long-horizon software development tasks that require hours of work. Program synthesis from scratch, as studied in ProgramBench, is particularly challenging when agents must infer behavior from usage documentation and an executable reference. Existing agents may lack concrete repair targets, lose track of explored and untested behaviors, and struggle to obtain trustworthy and timely progress feedback. We introduce , a framework that enables coding agents to manufacture and update an explicit progress signal from executable tests grounded in reference behavior. The framework preserves tests in a persistent corpus to provide reproducible repair targets and regression checks, while an obligation ledger guides further probing toward untested documented behaviors. A new test selection strategy controls redundancy and corpus growth through behavioral classes, separate quotas for divergent and regression tests, and explicit admission and corpus limits. When results are averaged equally across the four model configurations, achieves a mean hidden-test pass rate of 44.49%, compared with 35.58% for MSA-full, 32.63% for Claude Code, and 39.17% for SpecFirst. These correspond to relative improvements of 25.02%, 36.35%, and 13.58%, respectively. Ablation studies support the contributions of corpus selection and the obligation ledger. Retrospective evaluations of candidate snapshots further present positive correlations between corpus and hidden-test pass rates, each averaged across tasks. These results indicate that the manufactured signal tracks improvements in average held-out performance during program synthesis. The code is available at https://anonymous.4open.science/r/codemaps-4F5B/https://anonymous.4open.science/r/codemaps-4F5B/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.