acceptodds
Under review as a conference paper at ICLR 2027

BenchMutator : Agent Mutated Benchmarks

Abstract

Benchmark data contamination undermines the evaluation of large language models. Transforming the instances of an existing benchmark is a lower-cost alternative to constructing a new one, but current transformation methods have two limitations. First, most are task-specific and do not transfer across benchmarks. Second, their effectiveness is commonly inferred from the performance degradation they induce, although degradation can also result from invalid instances, increased difficulty, or variance across generation runs. We introduce BenchMutator, a task-adaptive multi-agent framework that calibrates transformation rules from benchmark examples and coordinates specialized agents to mutate inputs, update reference answers, and validate the resulting instances. We evaluate BenchMutator on six benchmarks spanning code synthesis, program repair, mathematical reasoning, and text-based question answering, using four model families. To control for contamination, we evaluate both base models and models deliberately fine-tuned on benchmark instances. We assess four properties, namely transformation diversity, human-assessed validity, consistency across generation runs, and model response to mutation. BenchMutator produces broader lexical and semantic variation than existing mutation approaches. A human audit reports 91.7% acceptance of transformed instances, with a near-zero mean shift in perceived difficulty. On GSM8K, BenchMutator yields lower run-to-run performance variation than the dynamic baselines Self-Evolving and DyVal. Mutation effects vary across tasks, generators, and model families, including for models fine-tuned on benchmark instances. These results indicate that performance degradation alone is not a reliable indicator of contamination mitigation, and that benchmark transformations should be assessed jointly on validity, diversity, and evaluation consistency. BenchMutator makes this assessment practical across task families.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.