acceptodds
Under review as a conference paper at ICLR 2027

Position Still Matters: Localizing and Controlling Latent Positional Effects in LLM Comparisons

Abstract

Large language models are increasingly used as evaluators, yet their judgments can depend on presentation order rather than only on relevant evidence. Evaluating both presentation orders can reveal inconsistent choices, but it does not show whether order still influences the decision when the final answer remains unchanged, where that influence enters the computation, or whether it can be reduced without weakening the comparison itself. We study these questions through a three-stage mechanistic framework. First, we decompose the model's decision margin into separate effects of comparison evidence and presentation position, revealing substantial positional shifts even in models whose choices are nearly invariant to swapping. Second, activation patching localizes this influence to late prompt states, while head-level interventions characterize how individual attention heads affect positional dependence and evidence sensitivity. Third, we construct interventions that reduce the positional effect under explicit constraints on comparison performance. On held-out numerical comparisons, the procedure removes 60–96% of the positional effect in five of six models, improving held-out balanced accuracy by up to 20 percentage points; in the remaining model, the same constraints permit only limited control. Reapplying the same procedure to country-population comparison and letter counting reduces the positional effect in all 12 model–task settings, by 21–100%, whereas the intervention selected on numerical comparison transfers inconsistently. These results show that behavioral order invariance can mask substantial latent positional influence, and that this influence can be causally localized and selectively controlled inside the model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.