acceptodds
Under review as a conference paper at ICLR 2027

Pay Once, Certify Every Round: Anytime-Valid Verdicts for Self-Improving Programs

Abstract

A program that improves itself repeats one round. It proposes a batch of candidates, edited copies of its own code, scores them on a fixed set of tasks, and keeps at most one. Each score costs a model call on one task, so the cost of a round is the number of program–task cells it reads. This paper asks how few cells a round can read, together with the decision rule that makes that small sample enough. We present BranchCert, a certificate that returns one of two verdicts for the round: Accept for a named candidate that beats the parent, the program the loop is running, or None Better when no candidate does. It returns Undecided when the budget runs out first. The certificate reads every candidate and the shared parent along one random order of the tasks. The unread tasks then form a uniformly random sample of the table, which lets the certificate stop as soon as the comparison is settled and bound its error at whatever time it stops. A second form spends the paid history of the loop. Earlier rounds scored other programs on the same tasks, those scores predict how the current batch will fare, and the certificate measures the residual the prediction leaves. On recorded rounds of a running self-improvement loop, on public evaluation tables, and on a paired re-execution of a recorded round, the certificate reaches the verdict of the complete table after a fraction of the reads. It agrees with the complete table in every run on the recorded loop, errs in at most 0.5% of runs on the public tables, and lowers its mean cost as the loop accumulates paid history.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.