acceptodds
Under review as a conference paper at ICLR 2027

Witness Before You Judge: Where Test-Time Program Selection Fails and How to Fix It

Abstract

Test-time scaling for code generation samples many candidate programs and spends LLM calls to pick one. When the pick is wrong, did the judge misread the evidence, or was the evidence never found? We study selection through witnesses, inputs on which two candidates behave differently. On a development pool of LiveCodeBench tasks, a pairwise LLM judge picks the correct program in all 60 comparisons in which it is shown a witness, yet 41 comparisons end without one. Time-outs account for 18 of them, and for 14 of these the first failing hidden test is longer than the inputs a judge is shown. We therefore separate evidence construction from judgement. WitnessFirst asks the LLM to propose objective witnesses, constraint-scale inputs from model-written generators, checked by a model-written validator, on which execution alone rejects programs that crash or time out; a judge then decides the disagreements that remain, reading the code when no witness is found. These eliminations are sound whenever the validator is and execution is reproducible, so each disagreement with a benchmark label is an invalid input, an unreproducible outcome, or a gap in the benchmark's tests. Across eight held-out candidate pools from LiveCodeBench and CodeContests+ and four judges, placing the objective stage in front of each of four selectors improves it with open-weight judges, by 3.2 to 5.7 points under the benchmarks' labels and 5.2 to 9.3 under audited labels; behind it, a six-call vote does at least as well as a 36-call one, and with Qwen3-32B the full WitnessFirst is the most accurate of the main selectors. In our pools, 1.9% of the benchmark-correct LiveCodeBench programs fail on legal inputs, spanning at least 13 of the 342 problems. With a frontier judge, reading the code alone is the most accurate selector, and objective witnesses still help the selectors that do not read code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.