EvoSampler: Utility-Guided Experience Sampling for Prompt Evolution
Abstract
Prompt evolution allows LLM agents to accumulate experience across rounds of revision and evaluation without updating model weights. Under a limited budget, an optimizer must choose which training instances to expose to the revision process, and this choice determines what experience can shape the next prompt. We argue for a relational view of training-instance value, contingent on whether the revisions it tends to induce address the failures the current prompt still exhibits. This view motivates an objective that scores candidates by their expected contribution to the validation gain of the post-revision prompt. We introduce EvoSampler, a utility-guided experience sampler that resolves this problem by learning from the history of past accept/reject outcomes and re-evaluating that history against the errors the current prompt still exhibits. EvoSampler estimates each training instance's reparability (its reliability in producing revisions that survive evaluation) and the validation failures its accepted revisions tend to resolve. These effects are captured by two complementary components: a Repairability Vector, which estimates each training instance's ability to induce successful revisions, and a Transfer Matrix, which characterizes the training-validation attribution between training instances and validation failures. An upper-confidence-bound rule further supplies exploration when evidence is sparse. Across five benchmarks spanning mathematical reasoning, interactive decision-making, and instruction following, EvoSampler achieves the highest average held-out score among the compared optimizers under both model configurations. With DeepSeek V4 Flash as the target LLM, it improves the HMMT and ALFWorld scores by 28.33 and 42 percentage points, respectively, over the unoptimized prompt. The implementation is available at [https://anonymous.4open.science/r/EvoSampler](https://anonymous.4open.science/r/EvoSampler).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.