What Does a Hallucination Score Tell Us? Auditing the Evidence from Prediction to Mechanistic Claims
Abstract
Internal representations offer a promising basis for detecting hallucinations in large language models and investigating how errors arise. Yet a score that predicts errors does not by itself establish which computation produced them or where an intervention should act. Connecting these claims requires tracing whether the observation, predictive support and intervention endpoint refer to the same target. We introduce a measurement-to-claim audit organized around three properties: risk observability, conditional quantitative readoutability and support identifiability. A separate intervention audit checks whether the output endpoint discriminates the target relation before testing selected-support effects against matched controls. Exact-prefix checks, candidate-conditioned readouts and support resampling establish the scope of predictive evidence; a controlled nonlinear calibration distinguishes known semantic-path supports from predictive proxies and response biases. Our main natural-model audit evaluates frozen supports on the same sources across four answer codebooks and both semantic mappings. Gemma retains internal predictive information across mappings, but output-endpoint discrimination and intervention specificity require separate evidence: even where both mapped endpoints discriminate the relation, a selected-support advantage over controls remains unestablished. A source-matched Qwen stress test stops earlier because none of its evaluated output endpoints reliably discriminates the relation, despite weaker internal readability. These results make the gap between prediction and mechanistic claims empirically traceable. The audit identifies which claim the evidence supports and which additional check is needed before a stronger interpretation is warranted.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.