acceptodds
Under review as a conference paper at ICLR 2027

Home-Away Evaluation: A Home-Field Framework for Pairwise LLM Evaluation

Abstract

Static benchmark accuracy summarizes how often a model succeeds, but obscures where its strengths differ from those of another model. We introduce Home-Away Eval, a pairwise evaluation framework that tests whether comparative strengths transfer to new questions. Each model generates questions from Home seeds that it answers correctly and its opponent answers incorrectly. Independent referees validate the questions and establish their answers; both players then answer the accepted questions on both fields. Only an exclusive correct answer earns a point. In our controlled benchmark-exposure experiments, fine-tuning improves static accuracy, but the fine-tuned checkpoints lose to their clean bases on the generated fields we evaluate. Training on the questions a comparison decides, by contrast, wins the match while leaving static accuracy slightly improved. An authorship study shows that each model answers its own questions more accurately than other authors' questions, which motivates reciprocal fields. The fields expose domain-dependent strengths, distinguish question posing from answering, and focus comparative evidence on questions the two models answer differently. Home-Away Eval complements static benchmarks with an explicit account of who challenges whom, on which questions, and with what evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.