acceptodds
Under review as a conference paper at ICLR 2027

Earlier Scores Count More: Causal Anchoring in Multi-Attribute LLM Evaluation

Abstract

LLM-as-a-judge protocols often score multiple criteria in a single response and aggregate them using an unweighted mean. Criterion order is known to shift scores, but correlation between earlier and later scores does not distinguish anchoring from shared dependence on item quality. We test for self-anchoring—the influence of a judge’s own earlier score tokens on its later scores—by intervening directly on generated score tokens. Holding the item, rubric, criterion order, and other preceding scores fixed, we set an earlier score above or below its baseline and measure the expected later score. Across five open-weight judges evaluated on teledermatology images and text, the mean controlled direct effect is positive for every judge and dataset. In both text settings, knocking out score-to-score attention removes the positive effect, while blocking intermediate positions does not. This dependence is concentrated in late decoder layers. Replacing earlier scores with a fixed reference value removes 97–98% of variance in aggregate scoring across criterion permutations. Measuring each position's total effect on the aggregate, with later scores generated sequentially, shows that the first output position carries roughly 5× its nominal weight in the smaller judges. Scoring one criterion per call improves agreement with dermatologists for most image criteria, but not for the aggregates, and it lowers correlation with human helpfulness on text. Equal nominal weights therefore do not imply equal criterion influence, and removing self-anchoring may not necessarily improve the validity of aggregate scoring.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.