acceptodds
Under review as a conference paper at ICLR 2027

Confident Hallucinations Live in the Blind Spot of Entropy

Abstract

Test-time methods that detect or reduce hallucinations through entropy cannot handle confident hallucinations. These wrong answers are as low-entropy as correct ones and persist under resampling. We study where their confidence comes from in context-grounded QA. Context ablation splits an answer's log-probability into a parametric prior and the evidence added by the context. Entropy bounds only their sum and is therefore blind to how the confidence is split. The split reveals two opposite sources of confident hallucination. Context-lure errors copy a wrong span from the context and have abnormally low parametric support. Prior-dominated errors ignore the context and have abnormally high parametric support. We trace both errors to retrieval that either copies a wrong span from the context or finds no matching span. Because benchmarks mix the two types, pooled grounding signals flip sign across datasets and cancel within them. To address this problem, we propose Grounding-Profile Contrastive Decoding (GPCD). GPCD is a label-free test-time method that corrects confident hallucinations without updating any weights. It detects answers that lean too far on one source by scoring them within context-support strata. It re-decodes each flagged answer with contrastive decoding in the direction opposite to its failure. It accepts a change only when resampled candidates agree on an answer that the context supports. On five context-grounded QA benchmarks, the stratified score detects confident hallucinations better than pooled signals on every dataset, and GPCD gives the best average accuracy while global contrastive decoding and self-consistency lose up to 7 EM points on average.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.