acceptodds
Under review as a conference paper at ICLR 2027

Convergent Elo: Online Pairwise Evaluation as Estimation

Abstract

Pairwise comparisons are increasingly relied upon to benchmark machine learning systems. When aggregated into a score for a fixed population, these comparisons pose an estimation problem: the reported ratings should settle as evidence accumulates. Fixed-gain Elo acts instead as a tracking algorithm: under noisy stationary comparisons, its ratings continue to fluctuate. We introduce Convergent Elo (CElo), which preserves Elo's two-participant online updates whilst decaying each participant's gain with respect to its own comparison count. We prove almost-sure convergence of the full rating vector under stationary outcome means and any fixed i.i.d. matchmaking distribution with connected support, without requiring global recentring. Further, we identify the limiting relative ratings: under Bradley–Terry realisability, CElo recovers latent strengths up to translation; under misspecification, it converges to the matchmaking-weighted KL-optimal Bradley–Terry approximation. Our analysis establishes a general convergence result for heterogeneous-gain stochastic approximation, deriving stability from coercivity on the identifiable subspace without assuming bounded iterates or imposing projection. In simulations of a stationary population calibrated from Arena Text data, CElo continues to reduce population-target error after fixed-gain Elo plateaus. A controlled ablation shows a clear finite-horizon advantage for participant-local over global gains under heterogeneous participation. Across four chronologically ordered applied datasets, CElo achieves lower prequential BCE than fixed-gain Elo on three, including text and vision model arenas, with no clear difference on the fourth.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.