acceptodds
Under review as a conference paper at ICLR 2027

On the Value of Iterative Repair in Code Generation under a Token Budget

Abstract

When a language model is asked to write a program and its first attempt fails the tests, the system spends further model calls to recover, and it does so in one of two ways: it either draws additional independent attempts from the same prompt, which is the resampling family behind pass@ evaluation and Best-of- selection, or it feeds the failing program and its test output back to the model for repair, which is the family behind self-debugging agents. Practice is split between the two and the published evidence points in both directions, since repair has been reported both as a decisive improvement and as a poor deal compared with simply drawing more samples. Which of the two is worth its price, and what decides the sign? In this paper, we first define the solve-within-budget curve of an inference policy under the provider's own token accounting and take its normalized area as the quantity to compare, so that a longer repair prompt is charged for what it costs. Then, we record two elementary facts that organize the measurement: the two families can differ only on episodes whose first attempt fails, so the benchmark-level contrast is a first-attempt failure rate times a conditional advantage, and a repair turn is charged for carrying the previous program, so repair cannot lead below the price of a first repair call. Finally, we measure both policy families on BigCodeBench, its hard subset and HumanEval+ across four model families, and we repeat the comparison with neither family capped in the number of calls it may buy. The experimental results show that repair leads resampling on demanding benchmarks once the budget affords a first repair call, that at the price of one further fresh call the two are indistinguishable, that the lead survives uncapping, that it rides on the execution feedback and persists when the graded tests are withheld from every prompt, and that the extra solutions carry a premium in tokens per solved task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.