acceptodds
Under review as a conference paper at ICLR 2027

Attention Sinks as Knowledge-Conflict Arbiters: Training-Free Gating-Preserving Routing for Contextual Faithfulness

Abstract

Large language models can generate unsupported answers even when the provided context contains sufficient evidence. Improving contextual faithfulness requires distinguishing answer-determining evidence from distracting information, rather than simply increasing attention to context. Attention sinks complicate this distinction: modifying their weights changes the attention budget available to content, while modifying their values changes the information they contribute. Our interventions show that both masking sink attention and replacing sink values with evidence-derived values tend to increase hallucination rather than improve grounding. These observations motivate a mechanistic account: sink gating is a decision variable the model actively modulates under answer-level conflict, and the sink value it weights carries a suppressed, prior-aligned direction. Building on this account, we propose AGPR, a training-free inference-time intervention that reads the model's own gating deviations as a conflict signal, requires layer-level consensus before acting, and applies a bounded correction along the direction the model has already chosen. Every routing correction is then projected onto a gating-neutral subspace via an exact normalization compensation, preserving the arbitration budget. Across four benchmarks, three model families, and multiple scales, AGPR reduces hallucination and improves grounded success while preserving general capability with negligible inference overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.