Answer Candidates Reshape Language Model Decisions: Layer-wise Dynamics and Causal State Interventions
Abstract
Multiple-choice benchmarks usually treat answer options as a passive interface for measuring what a language model knows. We show that answer candidates actively shape how models form answers. For each question, we construct paired inputs that preserve the question and correct answer while replacing only wrong options with unrelated alternatives from the same task. Across six Llama and Qwen models ranging from 3B to 14B parameters and fifteen tasks, we find that benchmark distractors delay (the top-ranked candidate) the last change in the model's favored option in 86 of 90 model-task settings, shorten the span over which the correct answer maintains a consistent lead (persistent correct-answer preference), and reduce accuracy by 20.2 percentage points. We then test whether these internal differences causally affect predictions. Specifically, we transfer the hidden state at the final token of the strongest wrong answer between paired inputs. This intervention improves accuracy by 8.78 percentage points and exceeds norm-matched random and cross-question controls. A separate 360-question panel further shows that reversing the transfer direction reverses the aggregate behavioral effect. These results show that answer candidates are not passive evaluation interfaces. Instead, they actively shape language model computation and causally influence final decisions. The code is provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.