acceptodds
Under review as a conference paper at ICLR 2027

Certified Reuse or Recalibration of LLM Judges after Output Shift

Abstract

An LLM judge is often calibrated on one model release and then reused after the evaluated outputs change. Its reliability map may continue to transfer, but it can also become stale when visible response properties change their relationship with quality. Recalibration is not automatically safer: generator replacement need not make a repair useful, and a small audit can favor a noisy repair. We study this as a finite-sample reuse-or-recalibration decision. Controlled style conflict provides a positive case, prespecified generator transitions test whether reuse remains adequate, and expert-scored long-form updates provide an external case. On a specified target window, a disjoint audit deploys a repair only when it certifies lower risk; reuse otherwise means only that no registered repair is resolved as beneficial at the current radius. Observable source–repair disagreement determines the paired-loss range, yielding safe near-oracle selection up to certification radii and an exact continuous minimax allocation. A fixed-repair binary-Brier lower bound matches the resulting worst-case label-complexity rate. Simulations verify finite-sample control; grouped replays show that certification suppresses held-out harm from validation-selected repairs while retaining power under drift.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.