acceptodds
Under review as a conference paper at ICLR 2027

The Tie-Breaking Degree of Freedom: Undocumented Protocol Choices Change Judge Leaderboard Winners

Abstract

Model leaderboards derived from LLM-as-a-judge evaluations or automated test suites frequently depend on arbitrary tie-breaking conventions to resolve exact score parity between competing systems. We investigate how sensitive leaderboard rankings are to these undocumented or minor protocol variations. By reanalyzing evaluation records across multiple benchmark suites, we isolate the specific impact of tie-breaking rules, rounding tolerances, and aggregation weights. Our findings reveal that while individual pairwise comparisons contain non-trivial tie frequencies, canonical leaderboard hierarchies remain remarkably robust to tie-handling modifications provided the overall score gaps between adjacent models exceed sample variance thresholds. We formalize a set of protocol-robustness checks and recommend that automated evaluation benchmarks explicitly report tie-resolution sensitivity metrics alongside point estimates to prevent fragile leaderboard inversions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.