acceptodds
Under review as a conference paper at ICLR 2027

Certified Early Rejection: Prompt Optimizers Can Skip a Third of Their Validation Rows

Abstract

Prompt optimizers such as GEPA, MIPROv2 and SIMBA spend most of their task-model calls scoring candidate prompts on every row of a validation set: in our GEPA runs these full evaluations take 97.9% of the calls, and four in five of the candidates scored end below the incumbent and are never returned. Stopping such an evaluation early is only safe if it is certified: under every possible outcome of the unseen rows, the chance of discarding a candidate that would have won must stay below a stated level. For a candidate evaluated on n rows in a random order, the largest expected saving of any admissible adaptive randomised score-blind rule is the value of an explicit linear programme, which we bracket in rational arithmetic; the same programme with a different error event governs screening a pool of discards. For each fixed recorded candidate stream, fresh level-alpha audits against the retained incumbent preserve the best presented recorded count with probability at least 1-alpha, without a multiplicity penalty. On 162 optimizer runs we executed (54 each of GEPA, SIMBA and MIPROv2; three task models, three benchmarks, 1,187 fully evaluated candidates), certified rejection at a 1% error budget, computed exactly on the runs' per-row evaluation records, skips 35.9%, 24.5% and 12.5% of the audited validation rows. A calibrated finite-population likelihood-ratio rule skips 37.2% of GEPA's rows, 97% of the certified per-candidate ceiling on the candidates below their incumbent. Run live inside the optimizers against paired controls, the certificate skips 28.7%, 11.9% and 10.5% of GEPA's, MIPROv2's and SIMBA's audited rows; on GEPA this is about two points below the exact replay expectation on the paired control streams, and it cuts task-model calls by 25.5%. The same theory explains the opposite regime: on a pre-registered cohort of generated programs, at the evaluated nomination budget and measured cost model, no admissible certificate screening the nominated pool could break even in expectation, even with perfect pool selection, and we compute how large, and how clean, the pool would have had to be.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.