acceptodds
Under review as a conference paper at ICLR 2027

HalluRefuse: Hallucination Mitigation with Knowledge Preserving Activation Steering

Abstract

Large language models (LLMs) perform well in question answering, reasoning, and text generation. However, LLMs may produce plausible but factually incorrect answers when they lack sufficient knowledge or are uncertain. For reliable deployment, LLMs should answer known questions correctly while abstaining from unknown questions. Prior work has shown that activation steering can steer LLMs toward refusal without additional training, but existing methods typically apply interventions to all questions without distinguishing between known and unknown ones. Through empirical analysis, we show that indiscriminate steering induces a trade-off between reducing hallucinations on unknown questions and preserving performance on known questions. We further find that known question representations exhibit a compact spectral structure, revealing a low-dimensional subspace that can be protected. Based on these observations, we propose HalluRefuse, an activation intervention method that mitigates hallucinations while preserving known representations. HalluRefuse estimates the subspace of known representations and applies the intervention to the orthogonal subspace. The derived closed-form intervention guarantees that the protected known representations are unaffected. Extensive experiments across multiple LLMs and question answering benchmarks show that HalluRefuse mitigates hallucinations on unknown questions while achieving the highest average known accuracy across all LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.