UntilItPasses: Benchmarking Repository Generation Against a Verifiable Goal
Abstract
Agentic coding capabilities have expanded from resolving individual issues to generating entire repositories. Existing repository benchmarks often evaluate agents in a single run and identify premature exit and unhandled edge cases as the leading failure patterns. However, real-world agentic coding involves feedback and multiple rounds of revision. To bridge this gap, we introduce UntilItPasses, an evaluation of 46 challenging repository-construction tasks that replaces the single submission with a minimal feedback loop. Each agent builds incrementally against a natural-language specification, commits its changes throughout development, and may make a feedback call at any point. Feedback reports only the fraction of hidden tests that pass, providing a milestone signal without revealing individual test cases. The agent continues working until it receives a perfect result or reaches a calibrated budget equivalent to six hours. Across twelve agent systems, 2.5% of runs exit on a perfect score. Our analysis of commit histories and feedback-call records traces agents’ progress through initial implementation and later repair, showing rapid early gains and persistent implementation gaps. Selected trajectories show agents using scalar feedback to broaden self-tests and revise mistaken assumptions shared by their code and tests. Model efficiency rankings also change with the required completion level. These findings highlight feedback-guided self-correction as a key capability to develop and evaluate in autonomous coding agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.