Depthwise Trajectory Signatures: Detecting Hallucination from Inside
Abstract
Large language models can produce fluent answers that are wrong, and most safeguards verify the answer only after it is complete. We take a different approach. Instead of checking what the model said, we monitor how it decided. Inside a transformer, a token prediction is built block by block, so the hidden state traces a trajectory across depth. We summarize this trajectory with a small set of geometric quantities, trajectory signatures, constructed to be invariant to the arbitrary choices that normally make depthwise measurements incomparable across layers, samples, and models. On these signatures we train a lightweight validator that detects hallucinations without modifying the base model. The same signatures localize the depth the validator finds anomalous, enabling a targeted correction at a single block during generation. Across 12 benchmarks and 5 LLMs, the validator achieves the best AUROC in 45 of 60 settings against baselines, and the refinement reduces the hallucination ratio by at least 10% relative in 29 of 45.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.