Taxonomy for Hallucination Localization and Attribution in LLMs
Abstract
While hallucination is a well-documented failure mode of LLMs, localizing and attributing its occurrence within generated text remains ambiguous. Existing span-level hallucination annotations often exhibit substantial disagreement because annotators select different textual spans for the same underlying semantic error. This ambiguity leads to noisy training labels, unreliable evaluation of hallucination detectors, and obscures where hallucinations emerge, knowledge that is essential for building effective mitigation strategies. In this work, we argue that this localization challenge originates due to the absence of a semantic model the defines the minimal unit of hallucination. To address this limitation, we introduce a diagnostic taxonomy for hallucination localization, grounded in Conceptual Dependency Theory (CDT), which categorizes hallucinations according to failures in five semantic components: , , , , and . A deterministic decision procedure guides annotators in identifying the minimal semantic component responsible for each hallucination, reducing ambiguity in span localization and semantic labelling. The taxonomy was developed through the analysis of 2,000 LLM hallucinations and validated in a held-out replication study in which five independent human annotators applied the frozen taxonomy and decision procedure to 300 previously unseen atomic claims drawn from four benchmark datasets. The taxonomy coupled with the decision tree achieved an average pairwise Cohen's kappa of 0.865, demonstrating that it reduces ambiguity in both hallucination span annotation and failure attribution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.