acceptodds
Under review as a conference paper at ICLR 2027

Agent as a Bench: What Do AI-Generated Tests Contribute to Evaluation?

Abstract

AI assistance accelerates code and research, increasing the need for reliable evaluation. Generated tests can expand this capacity, but construction success does not establish their value on other implementations. Our study identifies two reasons: detections overlap existing suites, and short packages retain different fault patterns across implementation populations. Agent as a Bench freezes generated packages to separate complementation, adding specification-supported detections beyond a complete suite, from compression, retaining its detections with fewer inputs. On HumanEval, the stronger suite already rejects 63 of 81 mismatch-selected targets. Generated tests expose one supported suite omission, while a fixed 120-target panel yields no confirmed additional detections. Compression nevertheless provides a measurable benefit: eight generated inputs retain five of six suite-rejected targets, compared with three for a random subset and four for a greedy subset, improving task-macro retention by 25 and 12.5 percentage points. These gains occur on one date task; development rankings differ. A year-range witness contributes 12.5 points against both subsets; the additional 12.5 points against Random8 involve a format-sensitive target. Uniform eight-input sampling from the fixed suite detects the year-range fault with probability 3.16%. The results establish added and retained detection as distinct sources of reuse value and show how fault-level evidence connects construction choices to that value.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.