When Aggregate AUROC Ranks the Shortcut: Trivial-Axis Headroom in Vision-Language Hallucination Detection
Abstract
Hallucination detectors for vision-language generators draw on four signal families (cross-modal alignment, self-consistency, an external classifier, token confidence) and are compared by aggregate AUROC. Proposing no detector, we introduce two diagnostics of what that comparison measures: headroom over a declared trivial axis—a label-free shortcut that never sees the image—and an exact decomposition of AUROC into the pairs that axis orders by itself and the rest. On chest X-ray report generation, our label rule marks asserted-but-unsupported claims hallucinated and denied-but-uncontradicted ones grounded, so reading silence as absence couples the label to claim polarity. A trivial axis reading only polarity reaches 0.84–0.87 AUROC on IU X-ray and MIMIC-CXR, leaving the best family a headroom of only 0.013–0.026 (one interval covering zero), whereas on COCO the best family clears a caption mention-count axis by 0.171–0.177. Because each axis lower-bounds what label-free scores can reach, these headrooms are upper bounds. Implementation choices move each family enough to reverse their ordering, yet at their best the three externally grounded ones converge on IU X-ray. That convergence is not agreement: on chest X-rays, polarity alone orders 70–74% of hallucinated–grounded pairs; cross-modal alignment gets all of them right yet is weakest or tied-weakest among claims denying a finding; and on MIMIC-CXR a classifier trained on neither corpus ranks last in aggregate (0.474, below chance) but first among those claims, and ordering claims by polarity and then by the classifier beats every family's best aggregate. Surface scores also depart from chance on three benchmarks labelled by others (label-free on POPE and M-HalDetect, label-informed on AMBER), even sign-flipped on POPE splits drawing negatives from frequent or co-occurring objects. Where labels are partly decidable from surface form, we recommend reporting headroom over a declared trivial axis, and stratified comparisons, alongside any ranking.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.