VERA: Verdict-Conditioned Reliability for Adaptive LLM Judges
Abstract
Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and poorly aligned with judgment correctness. We propose **VERA**, a *Verdict-conditioned Reliability Axis* that estimates reliability from hidden activations by distinguishing correct from incorrect judgments within each predicted-verdict group. Using **VERA** as a control signal, we develop a **VERA-guided periodic adaptation framework** that integrates reliability-ranked corrective updates, reliability-residual replay, and periodic refresh of the reliability directions. After VERA-guided adaptation on Chatbot Arena, 8B- and 14B-parameter judges outperform the strongest baseline on each of four held-out public benchmarks, with relative gains of up to **23.01%**. The framework also improves focal-class recall by up to **16.1%** relative to the strongest adaptive baselines on a separate proprietary temporal auditing task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.