DASH: Mitigating LLM Hallucinations via Dynamic Attention Steering in Hidden-Space for Contextual Faithfulness
Abstract
LLMs frequently hallucinate when parametric memory overrides the provided context. Existing contrastive decoding methods rely on static, input-agnostic retrieval-head masks and linear logit-space corrections, both of which limit faithfulness. We introduce DASH, which learns contextual faithfulness atop a frozen base LLM via a hierarchical Head Selection Controller that dynamically gates retrieval heads from per-token geometric features and a Nonlinear Steering Adapter that emits counterfactual residual corrections natively in the model's hidden space. We show that these corrections strictly contain affine contrastive logit steering as a one-parameter linear special case, and that the controller's optimal gate distribution admits a closed-form softmax characterization. Across single and multi-hop reasoning benchmarks, DASH establishes state-of-the-art groundedness on exact-match, F1, and ROUGE while keeping its generations fluent. We further provide DASH++, an entropy-triggered early-exit inference mode of the same model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.