acceptodds
Under review as a conference paper at ICLR 2027

Fails Before Repair Is Not Enough: Auditing the Validity of Executable Repair Benchmarks.

Abstract

Executable repair benchmarks are trusted because they run. A task counts as valid when the buggy artifact fails its tests and the reference fix passes them, and this two-sided check is close to the only validity evidence most benchmarks report. We show it misses three distinct failures, using a suite of 81 Verilog repair tasks and three instruments that apply to any executable repair benchmark. First, answer leakage: in 89.5% of our tasks the fix was written into the task inputs, and no task withheld it. Removing it costs a 1.5B-parameter model 16.9 points of pass@1 (95% CI [9.5, 24.7]) and four API models between -0.6 and 3.3 points, and with the answer withheld the share of tasks that separate two frontier models rises from 14.8% to 32.1% (paired +17.3 points, [7.4, 28.4]). Second, oracle weakness: our testbenches detect only 61.6% of injected single-site mutants, and four detect none. Applying the same mutation instrument to two benchmarks we did not build locates the cause in oracle design. RTLLM, whose testbenches are directed like ours, scores 52.9%; VerilogEval, whose testbenches compare the candidate against a reference cycle by cycle, scores 91.8%. Third, resolution: even with the answer withheld, 36 of 81 tasks are solved by all four API models on every sample. We also report that mining frontier-model failures into new tasks produced no novel tasks, with or without the answer available, and why. We release the suite, the harness and all three instruments.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.