When the Judge Changes the Verdict: Rank Instability and an Auditing Protocol for LLM-as-a-Judge
Abstract
LLM-as-a-judge reliability is typically treated as a model property, but we show it requires auditing along two orthogonal axes. Rank stability across sessions is not guaranteed: even with identical inputs, greedy decoding, and the same weights, a served judge can realize distinct deterministic scoring functions across sessions. This failure—operating as a score re-permutation across documents—is invisible to within-session retests and distributional checks, yet reverses downstream difference-in-differences estimates on up to dimensions. Calibration to a shared reference is pervasively weaker: across – open-weight judges, within-session agreement is near-perfect (), yet alignment to a production scale correlates at only – (median ). We propose two mechanism-agnostic protocols: a 20-document reference anchor detecting session deviations, and multi-session consensus recovering majority-mode estimates when the anomalous mode is minority. On 188k legal documents, the anchor classifies all 16 campaigns perfectly; consensus eliminates worst-case sign-flips when and reduces them three-fold at . The framework generalizes to new tasks or stacks by substituting the anchor; external validation of the reference mode against human annotations remains outstanding for this domain. Reliability is a property of a measurement campaign, not a model alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.