Reliable LLM-as-a-judge Evaluation via Transferable Calibration
Abstract
LLM-as-a-judge (LaaJ) has become a standard tool for scalable evaluation, but its reliability remains a central concern. Conformal inference (CI) can provide user-specified coverage guarantees for LaaJ outputs, but only under exchangeability. Shifts in evaluated models, topics, or response policies break this assumption and invalidate the coverage guarantee. Existing shift-aware CI methods theoretically suggest transferable calibration, but in LaaJ they face a covariate-nuisance bottleneck. We propose Robust Weighted Conformal Inference (RoWCI), which maps each example to a low-cardinality representation \(Z\), capturing coarse, transferable internal evaluation states. Through -state shift correction, RoWCI reuses source calibration data with human scores, without target-domain human scores. We further show that, under covariate shift, RoWCI provides a provable coverage degradation bound in the target domain, even when calibration is transferred without target labels. Empirically, across shifts in evaluated model, response policy, and topic, RoWCI achieves coverage closer to the nominal level with lower variance and competitive interval size than post-hoc LaaJ score calibration baselines. To our knowledge, RoWCI is the first approach to design shift-aware CI for reliable LaaJ evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.