When Can Historical Scores Be Reused After Judge Replacement?
Abstract
When an LLM judge is replaced, a pipeline that stores its scores must decide whether to keep the historical scores or pay to score every object again. We ask when reuse supports a stated downstream decision policy, and when establishing that is worth its cost. In three revision-pinned open-weight replacements, better discrimination and compatible decisions come apart: a Gemma replacement that separates clean from flawed proposals more sharply newly rejects 43 of 217 pool proposals, each of which its predecessor had accepted. For fixed objectwise policies, inverting the hypergeometric test bounds disagreement on the unaudited part of a finite pool, and the bound yields an audit-budget frontier that can be checked before any new judgment is bought. With 95 objects, no audit smaller than 33 can approve a 10% tolerance, and none smaller than 74 can approve 5%. Exact top-k membership, in contrast, is not identified by a partial audit. We replay the audit on saved judgments from 14 adjacent and 4 cumulative replacements served through commercial API gateways, where a model name is a routing label rather than a verified snapshot. At 10% tolerance with a 48-object audit, the best protocol saves only 2.62% of new-judge calls on average, and restarting a full batch after a failed audit costs more than rescoring outright.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.