Digits Don't Lie: Detecting LLM Hallucinations through Activation Digit Distributions
Abstract
In large language models, wrong answers look as fluent as right ones, so the text is weak evidence about correctness, but the activations carry more. Many activation-based detectors read that evidence from thousands to tens of thousands of neuron-aligned coordinates, each tied to one neuron of one model, so the evidence is large and has no fixed meaning across models. We introduce DigitTrace, which summarizes how values are distributed, not which neurons produce them. Leading digits do this by describing magnitudes on a common relative scale. We reduce six activation streams at every fourth layer to their leading-digit frequencies, Benford-reference statistics, and activation moments, accumulating these during generation into a trace under a thousand coordinates. Against seven published detectors on three models and eight benchmarks under one protocol, DigitTrace comes within 0.55 AUROC points of the best of them while using one tenth as many coordinates, and it shows particular strengths on mathematics and science. The compact trace grows more informative during generation: on seven matched cohorts it gains 10.37 points from early to late, while a raw-state control barely moves, so later tokens add evidence at fixed size. The trace can be inspected, not only scored: in most settings, correct and incorrect answers favor different leading digits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.