EvalOpt: Evolving Evaluators for Open-ended Tasks with Human Feedback
Abstract
Evaluating open-ended outputs requires task-dependent judgments across multiple aspects of quality that must be balanced in context. Existing proxy metrics and model-based evaluators do not consistently align with human preferences. Humans can often tell which of two outputs is better, but turning that judgment into an effective evaluation procedure is harder. Evaluator self-evolution offers a way to learn such procedures from human feedback, but the direction of improvement is often ambiguous. The same evaluation error can suggest several plausible corrections, while a useful direction may require multiple revisions before producing a measurable gain. An unsuccessful early attempt therefore provides limited evidence about whether a direction is ineffective or simply underdeveloped. In the default model-driven evolution setup we tested, unsuccessful early candidates were often followed by shifts to alternative directions. We present EvalOpt, a framework for self-evolving evaluators across open-ended tasks with fixed base-model weights. EvalOpt combines disagreement-based preference mining with depth-first evaluator evolution, explicitly maintaining an improvement direction across revisions until diagnostics support moving on or a rejection or resource limit is reached. Unsuccessful candidates guide further implementation with-out replacing the best evaluator retained on validation data. Across 19 domains, EvalOpt reaches 69.6% held-out decisive-pair accuracy, exceeding the initial Seed evaluator and GEPA, the strongest baseline, by 8.0 and 5.5 percentage points under matched feedback and search budgets. On four analysis domains, it exceeds fixed-patience continuation by 3.9 points, while preference mining adds 1.8 points over random sampling in existing-label replay at the same feedback budget. Independent blinded judgments favor EvalOpt-guided over Seed-guided outputs at tie-aware preference rates of 60.8% for critique-guided refinement and 59.1% for reward-guided training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.