RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
Abstract
As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes. LLM judges make this evaluation scalable, but the judges themselves must also be evaluated to ensure that their assessments are reliable. Benchmarks built around isolated responses do not test failures that depend on conversation history, and fixed benchmarks can provide less differentiation among increasingly capable models. Testing judges for such interactions requires multi-turn benchmarks that can be renewed as models and applications evolve. We introduce RankJudge, an end-to-end generator of document-grounded, multi-turn judge benchmarks. Its fully synthetic construction produces labeled conversation pairs through controlled generation and automated verification, without per-item human annotation. A judge-versus-problem formulation then jointly estimates judge ability and pair difficulty using the Bradley-Terry model. These estimates support difficulty-based benchmark curation and judge ranking under incomplete evaluation coverage. We instantiate RankJudge in machine learning, biomedicine, and finance, and evaluate 21 LLM judges, revealing substantial differences in judge ability across models. RankJudge thus provides a reusable construction and calibration pipeline, rather than only a fixed evaluation set.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.