Placebo Audits Quantify Audit-Specific Acceptance under Adaptive Reuse of Verifier Tests
Abstract
Coding agents are often given diagnostic feedback (failing inputs, expected and actual outputs) by the same small test suite that later decides whether their code is accepted. Holdout-reuse theory warns that this lets an agent fit the audit instead of the task. We make this measurable and show how to avoid it. A placebo audit re-evaluates each frozen output on a fresh audit drawn from the same conditional law as the real one. The gap between the two acceptance rates, audit-specific acceptance, equals the audit-specific part of false acceptance, is zero for any policy whose output is independent of the audit given the first attempt, is bounded by the information the feedback carries about the audit, and has a closed-form unbiased estimator. The same argument gives a remedy: diagnostics computed on a feedback set drawn independently of the acceptance audit provably add no audit-specific acceptance after one round. On 540 EvalPlus tasks, with four models from three families (8B to 32B), four audit sizes, up to 8 feedback rounds and eleven full runs, ten of them preregistered, the placebo meets its calibration prediction (i.i.d. resampling and binary feedback: +0.2 and +0.1 points after one round). One round of diagnostic feedback yields +2.9 [+1.8, +4.1] points of audit-specific acceptance for Qwen3-8B (four runs) and a positive estimate for every other model, including +2.6 [+0.8, +4.5] in a preregistered test at 32B; the point estimates assign roughly half of diagnostics' excess false acceptance to it. After eight retries, i.i.d. selection over nine candidates reaches a similar level. Independent feedback audits kept a fraction 0.74 [0.54, 0.95] of the diagnostic gain in useful acceptance, with audit-specific acceptance +0.1 [-1.0, +1.3]. On 444 LiveCodeBench problems, one-round audit-specific acceptance is +2.6 [+1.0, +4.3], and the preregistered precision contrast is inconclusive. We also report a preregistered replication failure of our own exploratory headline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.