acceptodds
Under review as a conference paper at ICLR 2027

From Prefill to Decoding: Rethinking Hallucination Detection Throughout LLM Inference

Abstract

Existing hallucination detectors for large language models (LLMs) typically focus on either the prefill or decoding stage. However, hallucination signals across these two stages have largely been studied in isolation. In this work, we reveal that prefill and decoding provide complementary signals for hallucination detection. Each stage captures distinct failure patterns that are often missed by the other. Motivated by these findings, we propose a unified hallucination detection framework that exploits the complementary signals throughout LLM inference. Our framework forms a prefill–decode detection cascade: it first adaptively probes intermediate layers during prefill, and then selectively verifies the remaining cases at critical semantic units as decoding unfolds. This cross-stage design allows the two stages to compensate for each other's detection blind spots, yielding stronger detection within a single decoding pass. It also reduces inference cost by using prefill to filter easy cases before expensive decode-stage verification. Across two open-source LLMs and three QA benchmarks, our method consistently outperforms all probe-based single-stage baselines while incurring substantially lower latency than decode-only detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.