The Tie-Breaking Degree of Freedom: Undocumented Protocol Choices Change Judge Leaderboard Winners
Abstract
Model leaderboards derived from LLM-as-a-judge evaluations or automated test suites frequently depend on arbitrary tie-breaking conventions to resolve exact score parity between competing systems. We investigate how sensitive leaderboard rankings are to these undocumented or minor protocol variations. By reanalyzing evaluation records across multiple benchmark suites, we isolate the specific impact of tie-breaking rules, rounding tolerances, and aggregation weights. Our findings reveal that while individual pairwise comparisons contain non-trivial tie frequencies, canonical leaderboard hierarchies remain remarkably robust to tie-handling modifications provided the overall score gaps between adjacent models exceed sample variance thresholds. We formalize a set of protocol-robustness checks and recommend that automated evaluation benchmarks explicitly report tie-resolution sensitivity metrics alongside point estimates to prevent fragile leaderboard inversions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.