acceptodds
Under review as a conference paper at ICLR 2027

A CaLLM Judge is a Better Judge: Mitigating Bias in LLM-as-a-judge

Abstract

Large language models (LLMs) are increasingly used as judges to predict which of two candidate answers a human evaluator would prefer, providing a scalable alternative when human evaluation is too costly. Unfortunately, LLM judgments can exhibit systematic biases. Prior work has documented position, verbosity or model family of the candidate answers influencing the judge decision rather than their quality, while additional biases may remain undocumented. Multicalibration provides a natural framework to address this problem: each bias corresponds to miscalibration on an identifiable group, e.g., position bias manifests as overconfidence in the answer shown first. The challenge is to fix miscalibration simultaneously across many, potentially unknown groups. To this end, we introduce CaLLM, a post hoc framework that multicalibrates an LLM judge, thereby bounding all such biases simultaneously. We prove that CaLLM bounds the bias of every group whose error structure is captured by an available representation, and that augmenting it with explicit bias groups further reduces the residual bias. We then derive theoretical conditions under which CaLLM improves the judge's accuracy, rather than trading it for calibration. Finally, we evaluate CaLLM across pairwise preference benchmarks, showing that it reduces the multicalibration error on both specified and unspecified bias groups, while improving predictive performance of the judge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.