acceptodds
Under review as a conference paper at ICLR 2027

Selecting Tests for Generated Code: Failure Coverage and Transfer Limits

Abstract

Test-time code generation is often improved by sampling several programs and selecting one that passes a limited set of tests. However, distinct test inputs can expose the same errors and miss failures specific to the target generator. We define certificate identifiability and distinguish identity-orthogonal evidence from mode-orthogonal evidence, which provides additional information about correctness. We then introduce transfer-aware Orthogonal Evidence Acquisition (OEA), which allocates a fixed test budget using naturally occurring proxy failures and structural support cases, and executes the selected evidence in fail-fast order. An impossibility result shows that complete coverage of a surrogate failure bank gives no distribution-free guarantee for natural target errors. We bound missed failure mass in terms of proxy coverage and distribution shift, and show why no ordering can detect failures unsupported by the test pool. On an official 96-task MBPP+ confirmation, support-hybrid budget-eight selection agrees with exhaustive-design candidate ranks on all 96 tasks while reducing online design calls by 83.8%. A prespecified ClassEval confirmation does not reproduce a holdout-only risk improvement. These results show that a small test subset can match exhaustive use of the design pool on MBPP+ under the frozen task-level protocol, but the benefit depends on which target errors the available tests can detect. The observed saving is an online design-call saving; it is not an end-to-end latency or compute claim. A support-matched failure-frequency baseline selects the same candidates with no greater online cost, so these results do not establish an additional benefit from proxy-mode diversification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.