Quantiles, Not Points: Streaming Quantile Matching for Ordinal-Aware LLM-as-a-Judge Calibration
Abstract
LLM-as-a-judge scores are often distributionally shifted from human ratings on any given evaluation task, and post-hoc calibration (mapping judge scores to human scores in a small labeled anchor set) has become a popular method to correct for this shift. We present streaming quantile matching (SQM), a method previously unexplored for LLM-as-a-judge calibration that wins outright over every other method by directly computing the optimal-transport map between the judge and human score distributions. We develop a novel framework for evaluating post-hoc calibration performance and within it we show that unlike point-wise calibrations, SQM respects the ordinality of the human anchor set and offers superior performance regardless of model or data budget. We also show that in general, post-hoc calibration narrows the raw error gap between smaller and larger judge models, which opens the possibility of more efficient LLM judges.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.