Planning to Explore: Curiosity-Driven Exploration for LLM Test Generation
Abstract
The use of LLMs for code generation has naturally extended to code testing and evaluation. As codebases grow in size and complexity, so does the need for automated test generation. Current approaches for LLM-based test generation rely on strategies that maximize immediate coverage gain, a greedy approach that plateaus on code where reaching deep branches requires setup steps that individually yield zero new coverage. Drawing on principles of Bayesian exploration, we treat the program's branch structure as an unknown environment and maintain an exploration state built from execution, represented by a coverage map and a per-function posterior over which parts of the program the agent can reach. Our method, CovQValue, uses this posterior to target candidate test plans at unexplored code, asks the LLM which functions each plan will execute and enable, and selects the plan with the highest estimated Q-value. We show that the major gains come from using the exploration state to decide what to explore next. Our method outperforms greedy selection on TestGenEval Lite, achieving 103-150% higher branch coverage and winning on 83-88% of targets. In addition, we build a benchmark for iterative test generation, RepoExploreBench, where CovQValue achieves 129-157% higher branch coverage than greedy selection across three LLMs, winning on 94-97% of targets. These results show the potential of curiosity-driven exploration for LLM agents, enabling more effective discovery with sequential interaction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.