acceptodds
Under review as a conference paper at ICLR 2027

GRA: Grounding Ratio Analysis for Prompt-Anchored Update Trajectories in Hallucination Detection

Abstract

Large language models (LLMs) can generate fluent but unsupported or factually incorrect responses, undermining their reliability in real-world deployment. Existing white-box hallucination detectors often focus on hidden-state snapshots or aggregate internal representations, leaving the generation-time update process underexplored. In this paper, we shift the focus from static hidden-state analysis to update anchoring and propose Grounding Ratio Analysis (GRA), a white-box hallucination detector that measures prompt anchoring in generation-time residual updates. For each generated token and Transformer layer, GRA measures whether a residual-stream update remains anchored to the prompt region or shifts toward the generated-prefix region, forming a token-layer grounding trajectory. We summarize this trajectory with mean-trend pooling and use a lightweight probe for response-level reliability estimation. Experiments on five open-weight LLMs and five question-answering and hallucination detection datasets show that GRA achieves the highest average AUROC among evaluated methods on every evaluated model. Further ablations show that temporal trajectory information is important, prompt-anchoring signals appear across different update sources, and raw GRA trajectory statistics are already directly separable. These results suggest that prompt-anchored update dynamics provide an effective and interpretable signal for hallucination detection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.