Read the Evidence Before Seeing the Answer: Hypothesis-Independent Evidence Reading for Detecting Confident Hallucinations
Abstract
Evidence-based fact verification normally asks a leading question: the candidate answer is placed in the verifier's context and the model is asked whether the evidence supports it. We study the confident, self-consistent errors on which such verification is most needed and least trusted. Auditing 9,762 of them across eight benchmarks shows a benchmark-dependent mixture of genuine hallucinations, ambiguous questions and wrong gold labels; on the genuine part the best internal signal we tried reaches 0.69 AUROC, versus 0.90 or more once the same model reads a passage, so evidence is needed. Candidate-conditioned verification has a measurable cost: holding question, evidence and verifier fixed and varying only whether the candidate is the verifier's own output, a model judges its own wrong answer more leniently than another model's equally wrong answer (19,718 matched pairs over five benchmarks, ), increasingly so with its confidence, and not at all when neither candidate is its own. We therefore make the evidence-interpretation stage hypothesis-independent: a reader answers the question from the evidence before seeing the candidate, , and an external matcher compares with the candidate, while the calibrated candidate-conditioned score is retained as a base term. This improves detection on every benchmark tested ( AUROC with benchmarks as units, ), by a similar margin whether or not the verifier produced the answer: what helps is reading the evidence first, not the removal of the ownership bias that motivated it. Routing only the questions that internal signals cannot separate matches the full-retrieval residual error rate at 39% retrieval coverage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.