BenchMutator : Agent Mutated Benchmarks
Abstract
Benchmark data contamination undermines the evaluation of large language models. Transforming the instances of an existing benchmark is a lower-cost alternative to constructing a new one, but current transformation methods have two limitations. First, most are task-specific and do not transfer across benchmarks. Second, their effectiveness is commonly inferred from the performance degradation they induce, although degradation can also result from invalid instances, increased difficulty, or variance across generation runs. We introduce BenchMutator, a task-adaptive multi-agent framework that calibrates transformation rules from benchmark examples and coordinates specialized agents to mutate inputs, update reference answers, and validate the resulting instances. We evaluate BenchMutator on six benchmarks spanning code synthesis, program repair, mathematical reasoning, and text-based question answering, using four model families. To control for contamination, we evaluate both base models and models deliberately fine-tuned on benchmark instances. We assess four properties, namely transformation diversity, human-assessed validity, consistency across generation runs, and model response to mutation. BenchMutator produces broader lexical and semantic variation than existing mutation approaches. A human audit reports 91.7% acceptance of transformed instances, with a near-zero mean shift in perceived difficulty. On GSM8K, BenchMutator yields lower run-to-run performance variation than the dynamic baselines Self-Evolving and DyVal. Mutation effects vary across tasks, generators, and model families, including for models fine-tuned on benchmark instances. These results indicate that performance degradation alone is not a reliable indicator of contamination mitigation, and that benchmark transformations should be assessed jointly on validity, diversity, and evaluation consistency. BenchMutator makes this assessment practical across task families.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.