acceptodds
Under review as a conference paper at ICLR 2027

Heads Take Sides: Imbalanced Static Voting in Unfair LLM Judges

Abstract

LLM judges can encode correctness-related information internally, yet how this information is converted into a final verdict remains unclear. By decomposing the verdict-logit margin into attention-head contributions, we identify a predominant head-level mechanism of **static voting** in LLM judges: opposing head groups consistently favor acceptance or rejection, while input-dependent voting strengths determine their competition. Although their voting directions remain fixed, their voting strengths vary with solution quality—acceptance heads weaken and rejection heads strengthen as errors accumulate—enabling input-dependent judgments through competition. This static voting structure recurs across model architectures and verdict formats, while targeted ablations confirm that the two camps causally push judgments in opposite directions. This account suggests that reliable correction should adapt the balance between the two camps rather than suppress either one. We therefore introduce **Voting Calibration**, which uses an internal correctness signal to modulate their voting strengths while preserving their specialized directions. Relying on a small reference set, Voting Calibration improves overall accuracy and true-negative rate by averages of 8.3% and 14.6%, respectively, across four models on -MATH benchmark.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.