DeCoS: Decoding by Contrasting Steered Features Reduces Hallucinations in Large Language Models
Abstract
Large language models (LLMs) can generate fluent but factually incorrect content, limiting their reliability in knowledge-intensive applications. We introduce Decoding by Contrasting Steered Features (DeCoS), an inference-time framework that mitigates hallucinations without external knowledge or model fine-tuning. Rather than contrasting dense layer or attention-head representations, DeCoS uses sparse autoencoders (SAEs) to identify hallucination-correlated latent features through prompt separation. At each decoding step, it constructs a strong-hallucination counterfactual and a weakly suppressed or unmodified reference branch, then contrasts their next-token distributions within an adaptive plausible-token set. We evaluate DeCoS on multiple-choice, open-ended generation, and chain-of-thought reasoning tasks across seven instruction-tuned models from the Gemma-2, Gemma-3, and Qwen3 families. DeCoS achieves the highest open-ended truthfulness–informativeness point estimate and the lowest rejection rate among the evaluated methods on all seven models. All 28 pairwise margins over four competing decoders are positive, with 16 nominally significant at the 5% level; the seven comparisons with the strongest per-model competitor are also positive but nonsignificant. DeCoS further achieves the best or tied-best score in all 14 model–reasoning-benchmark comparisons. Results on four pretrained models and complementary sensitivity analyses provide additional evidence for the method's robustness across the evaluated settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.