Token-Level Evidential Support Is Linearly Decodable from the Residual Stream
Abstract
We replace hallucination detection with a relaxation that is decidable by construction: for each token a model generates, was the evidence for it present in the context, and does the model's internal state reflect that? Each item is a two-part question about a 10-K filing, with the supporting sentence for one part deleted and replaced by a length-matched sentence from another filing, alternating which part is affected. Because we delete it ourselves the label is exact rather than annotated, and because both parts are answered in one generation the two classes share prompt, document, question and output. A linear probe on mid-network activations ranks the supported part's tokens above the withheld part's at AUC, bootstrap interval over answers, against a bar of fixed before the data was collected. The probe also localizes, which a single score for a whole response cannot do. Thresholding its per-token score recovers the supported part's tokens at a token-level intersection over union of , against for next-token entropy read at the same position. At a matched scoring unit the probe leads entropy by of AUC per token and per part. Three pre-declared controls come out at chance, as does a randomly initialized placebo, and two positive controls come out high, which is what makes the battery falsifiable. One leak is real: a bag-of-words classifier on the answer text alone separates the conditions across answers at , so the answer-level gate is benchmarked against that rather than against chance. Within one answer the contrast is not recoverable from which tokens were generated, and that is the control the main result rests on. The probe reads support and not accuracy: answers whose evidence was present but whose value was wrong are indistinguishable from correctly answered supported ones, while entropy places those same answers at its own midpoint. This matters most on long documents, where the model stops flagging the problem itself. Padding the document from to tokens with the question and the removed evidence held fixed, the rate at which the model declines to answer falls from to . The depth at which the signal appears does not move with length, on a grid of four probed layers, refuting a hypothesis we had registered. What the probe provides is a ranking within an answer, not a calibrated absolute score, and a pre-registered test of locality failed. Every result here comes from a single B-parameter model, with the headline replicated on one other.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.