acceptodds
Under review as a conference paper at ICLR 2027

VERA: Verdict-Conditioned Reliability for Adaptive LLM Judges

Abstract

Accurately estimating judgment reliability is a central challenge in adapting LLM judges to newly verified feedback while preserving previously learned behavior. However, existing approaches often rely on output-level confidence, which can be overconfident and poorly aligned with judgment correctness. We propose **VERA**, a *Verdict-conditioned Reliability Axis* that estimates reliability from hidden activations by distinguishing correct from incorrect judgments within each predicted-verdict group. Using **VERA** as a control signal, we develop a **VERA-guided periodic adaptation framework** that integrates reliability-ranked corrective updates, reliability-residual replay, and periodic refresh of the reliability directions. After VERA-guided adaptation on Chatbot Arena, 8B- and 14B-parameter judges outperform the strongest baseline on each of four held-out public benchmarks, with relative gains of up to **23.01%**. The framework also improves focal-class recall by up to **16.1%** relative to the strongest adaptive baselines on a separate proprietary temporal auditing task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.