How LLMs Are Persuaded: A Few Attention Heads, Rerouted
Abstract
Large language models can answer a factual question correctly, yet switch to an incorrect answer when exposed to persuasive context. Why this happens inside the model remains unclear. In this work, we study this question through causal interventions in controlled multiple-choice tasks and identify that persuasion mainly acts through a small set of mid-layer attention heads, which we call *decision heads*. These heads do not compute the answer themselves. Instead, they copy the option they attend to, so changing their attention can directly change the model's answer. We find that this attention is controlled by a one-dimensional *option-routing feature* in the option-token representations. Persuasive text changes this feature, causing the decision heads to attend to the persuasion target instead of the correct option. We further trace the routing feature to shallower attention heads that read persuasive keywords from the input. Together, these results reveal a causal pathway: shallow heads read persuasive keywords and write the routing feature, which redirects the decision heads’ attention toward the persuasion target; the decision heads then copy the selected option, leading to the wrong answer. We verify each step through intervention and find the same mechanism across four open-source model families and in a source-selection task derived from Generative Engine Optimization (GEO). Finally, based on this mechanism, we introduce a low-rank update to the decision heads' key projections to reduce their sensitivity to persuasion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.