Topological Hallucination Detectors on Decoder Attention Reduce to First-Order Statistics
Abstract
Persistent homology (PH) of attention graphs underlies a family of hallucination detectors, including TOHA. We show that in decoder-only models these topological features reduce to first-order attention statistics — quantities read off single rows or columns of the attention matrix, such as a row maximum or one token's attention to another — either exactly or up to one computable number. For any symmetric dissimilarity and vertex we define the coning defect ; classical cone and stability arguments then place the whole Vietoris–Rips barcode within bottleneck distance of a "star diagram" read off the distances to , so every stable vectorization is a function of one column of the attention matrix, up to . The same argument applies to TOHA: its prompt–response topological divergence equals one minus the response tokens' mean largest prompt attention whenever a prompt-level coning defect vanishes, and lies within that defect otherwise. Across 16 settings (7 open models of 1.1B–7.62B; TruthfulQA, HaluEval, and on-policy TriviaQA) the identity holds without violation on 20,580,160 per-head graphs and is exact on 57–88% of them, including in two models with no first-token sink, where the statistic is the largest prompt attention rather than sink attention. Adding TOHA's score to that first-order counterpart changes AUC by at most 0.007, and supervised probes on the two are equivalent in every setting. On 838,356 layer graphs there are no bound violations; in every model with a first-token sink, 36–98% of graphs are exactly coned with empty , and length-normalized 0D persistence keeps a median per-layer rank correlation of at least 0.98 with sink attention. For detection, 0D PH, sink-deflated PH and the per-head coning defect add at most 0.009 AUC over cheap non-topological probes, and TOHA at most 0.013. We release the extractor, the evaluator, and per-layer features for a subset of settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.