acceptodds
Under review as a conference paper at ICLR 2027

BEHAVIORAL TEST DIVERSITY: QUANTIFYING TEST SUITE DIVERSITY THROUGH AGENTIC CODE REPAIR

Abstract

As coding agents make it easier to generate code at scale, ensuring that the resulting code aligns with human intent becomes a central challenge. Testing is widely used to check whether code behaves as intended, with tests serving as executable specifications. The increasing effort required to construct tests has motivated growing use of LLM based test generation, which requires effective objectives to guide test suite construction. Existing objectives do not directly quantify the diversity of behavioral constraints specified by a test suite. High performance under these metrics can therefore coexist with superfluous redundancy of some behaviors and omissions of others, leaving gaps in the testing surface. We introduce Behavioral Test Diversity (BTD), an agent conditioned metric that measures this diversity and provides an objective for test suite optimization. BTD quantifies behavioral overlap by observing whether code repairs guided by one test also satisfy others, without needing predefined behavior categories. The resulting score serves as an objective for preserving behavioral diversity during test suite pruning and expansion. Our experiments show that BTD yields stable scores and preserves relative test suite rankings. Across our evaluations on ProgramBench, vLLM, and ComputeEval, BTD guided pruning reduces the number of tests supplied as additional executable specifications by 61 to 90% while retaining sufficient behavioral guidance for full suite recovery. We also show that BTD guided expansion covers previously unchecked behavior that baseline methods miss. Together, these studies support BTD’s validity and practical utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.