Small but Mighty: Quality–Diversity Evolutionary Search for Size-Optimized Unit Test Generation
Abstract
Reliable coding agents require verification signals that are not only strong but cheap enough to execute repeatedly within agentic loops. Yet existing unit-test generation methods optimize coverage or mutation score while leaving suite size uncontrolled, often producing increasingly redundant tests. We formulate size-optimized unit test generation as a quality-diversity problem: maximize behavioral coverage while discovering compact suites along the quality–size frontier. We introduce TeSearch, an inference-time evolutionary search that partitions test suites into size niches via MAP-Elites and maintains a Pareto frontier over coverage and mutation score within each niche. A size-annealing UCB allocates generation budget across niches, prioritizing quality early and progressively favoring smaller suites once quality is secured. Crucially, quantitative control through an LLM is itself stochastic: requesting k tests does not reliably produce k tests. TeSearch therefore models the LLM as a noisy actuator through an online landing kernel and compensates for its size-instruction bias before generation. Across 129 TestGenEval modules and three frontier LLMs, TeSearch outperforms conventional inference-time scaling and coding agents, achieving higher hypervolume with substantially smaller suites. The optimized suites further improve test-guided cross-version migration and provide fast, effective acceptance checks for self-evolving agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.