Decision Geometry of Pairwise LLM Evaluation
Abstract
Changing the labels around fixed answers can change an LLM judge's decisions. Interpreting the resulting score gain requires understanding which choices are made and how they combine across candidate orders. Our label-intervention studies show that similar correctness gains can arise from different changes in decisive coverage and alignment with the correct answer, with different implications as the cost of errors changes. We trace these effects from paired verdicts to their decision geometry. An invertible decomposition gives an exact attribution of correctness changes to coverage and signed alignment. A threshold model then explains how position preference and tying interact: a tie can reveal a preference that opposing votes would cancel, while a higher threshold can remove it again. The model yields a non-monotone selection boundary and a sign-invariance prediction under shared evidence. Response fits favor label-dependent position and tie parameters, and paired diagnostics distinguish observed cross-label patterns from both deterministic sign invariance and the higher reversal frequencies predicted by independent completion. Sharp coupling bounds characterize the paired energies compatible with each fitted marginal response law. This analysis follows an interface change through selection, aggregation and joint behavior to explain what changes behind its reported score. An anonymized code snapshot is available at https://anonymous.4open.science/r/dgpllme/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.