acceptodds
Under review as a conference paper at ICLR 2027

Answers as Probes: Mitigating LLM Hallucinations with Source-Aware Inference

Abstract

Large language models that answer from a given passage still produce answers the passage does not support. We find that many of these errors are failures of reading, not of knowledge: if Llama-2-7B-chat reads each MRQA prompt in four slightly different ways and the best of its four answers is kept, macro F1 rises from 67.0 to 74.7. Decoding-time methods recover little of this gap, for two reasons. They build the alternative reading by deleting the context, which also changes how the question is encoded, and they commit to one reading token by token, before the answer that would reveal whether the reading was right exists. We propose , a training-free decoder that addresses both. It intervenes only on what the answer may attend to, which leaves every state before the answer intact and lets several readings share one prefill and decode in one batch. Each reading proposes a complete answer; each answer then probes its source, and the unmodified model selects the one with the highest native belief plus source support: exact containment for text sources, an interventional likelihood ratio otherwise. On all 62,316 MRQA questions, improves every domain for each of five models, gaining 3.0–4.2 F1 over greedy decoding and 2.1–2.3 F1 over the strongest published decoder at greedy cost, with no tuned hyperparameters. It also improves TruthfulQA MC1 for all five models tested. Used as an instrument, it shows that hallucinated answers do not read the passage less than correct ones, but read it through shallow layers and punctuation anchors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.