acceptodds
Under review as a conference paper at ICLR 2027

ProbeTok: Real-Time Hallucination Detection in RAG with One Shared Probe Token

Abstract

Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination: a response may still contain sentences that the retrieved documents do not support. Such a sentence should be detected in real time since LLMs tend to build on their earlier mistakes. However, most detectors check the complete response and therefore run only after it has been generated, and most of those that can check each sentence in real time need a separate model to read the documents. In this paper, we rethink real-time detection from the perspective of generation and find that, at the moment a sentence is generated, the generator has already read everything a detector needs and holds it in its cache, which makes it a natural basis for a real-time detector. However, its hidden states are computed to generate the next token, not to detect hallucination. Intuitively, asking the generator whether the sentence is supported by the documents would switch it from generation to detection, but asking after every sentence is costly. We therefore propose ProbeTok, which uses a single learned token, the probe token, instead of a handwritten question, and in our experiments the learned token is on average more accurate than the handwritten one. After each sentence is generated, ProbeTok feeds the probe token to the frozen generator. The embedding of the probe token is learned so that the generator computes a hidden state for detection, from which a linear head reads the hallucination score. ProbeTok then removes the probe token and restores the cache, so the generator continues the response as if no check had happened. On the RAGTruth and TRIVIA+ datasets with responses from four of the latest LLMs, ProbeTok achieves the highest average sentence-level AUROC compared with state-of-the-art detectors, and adds only 4.2% to the generation time. Moreover, a strong open LLM can use our method to check the responses of the other generators without access to their hidden states. In our experiments the best such detector is on average only 0.03 AUROC below ProbeTok running on the generator itself.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.