HPSA: Hallucination Detection from a Single Hidden State, Without Training or Layer Search
Abstract
Large language models can produce fluent but unsupported statements, yet many detectors require labelled in-domain data, additional generations, or a trained probe. We study a simpler alternative: a single hidden state from the answer-generation pass, robust coordinate scaling, and a closed-form first-moment direction. The resulting detector, HPSA, uses no fitted classifier, extra sampling, or target labels to produce its score. Across four question-answering benchmarks and three dense decoder models, it reaches a mean out-of-fold ROC-AUC of over twelve model–benchmark cells and is the best detector without labels from the scored rows in seven cells. A geometric analysis finds a consistent location shift between the two label populations. Finally, a direction estimated from other benchmarks retains of the target's own label-free AUC on average. These results support a narrow conclusion: a fixed hidden-state statistic ranks answers usefully at the cost of one dot product.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.