acceptodds
Under review as a conference paper at ICLR 2027

What Can Be Corrected in LLM Judge Ratings? Identifiability, Measurement, and Targeted Adjustment

Abstract

Correcting ratings from LLM judges requires knowing whether associations with features such as verbosity, lexical overlap, or displayed position reflect bias or item quality. Ratings alone can identify differences in these associations across judges, but cannot separate an association shared by the panel from item quality. We formalize this limit with a hierarchical Bayesian measurement model that represents each ordinal rating in terms of latent item quality, judge leniency, and feature associations. We use this model as a diagnostic audit: human labels, controlled counterfactuals, or another external anchor are required to identify a panel-wide association. We therefore vary unidentified panel-wide slopes in sensitivity analyses. Targeted adjustment is a diagnostic control that removes only identifiable judge-specific deviations. On TREC web relevance and SummEval summarization, targeted adjustment has higher Spearman rank-correlation point estimates with human labels than full adjustment, although panel averaging matches or outperforms it. As a practical implementation layer, a 250,000-parameter amortized network matches sampling-based masked-rating reconstruction error (0.302 versus 0.304) while reducing per-dataset inference time from about 240 seconds to 1.3 ms. We recommend selecting a scoring rule through human-label validation and removing a panel-shared association only when an external anchor identifies the corresponding effect.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.