Trust the Stream: Online Confidence Calibration for Selective LLM Judges
Abstract
Large language models (LLMs) are increasingly used as automated judges, motivating selective evaluation methods that certify human agreement by accepting only sufficiently confident predictions. However, with limited calibration data, noisy estimates of the confidence distribution can destabilize the certified threshold and degrade the coverage–agreement trade-off. We introduce Online Confidence Calibration (OCC), which continuously refines the empirical confidence distribution using confidence scores observed during deployment. OCC maps confidence scores to their empirical percentiles, re-scores the fixed calibration set, and re-certifies the acceptance threshold without requiring additional human labels. We prove that the resulting threshold converges to its population counterpart, achieving finite-time exact recovery with high probability and eventual almost-sure lock-in, while preserving the human-agreement guarantee at every round. Across six benchmarks, eleven LLM judges, and four confidence functions, OCC consistently improves threshold stability and calibration quality, certifies stronger agreement at matched coverage than label-free baselines, approaches the population-CDF oracle.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.