Fair Protocol Reveals Collapse of Medical Hallucination Detectors-and Two Detectors That Survive It
Abstract
State-of-the-art hallucination detectors for medical VQA fail to reliably measure whether answers are grounded in the image. We show that reported progress is largely an artifact of benchmark leakage: detectors exploit surface statistics like answer length, formatting, and sampling dispersion that correlate with hallucination only within the benchmarks used for tuning. We introduce a fair evaluation protocol that strips away these shortcuts via length- and type-matched negatives, image-level splits, and language-leakage controls. Under this protocol, all prior methods collapse to near-chance performance. VASE, SE, RadFlag, and a frozen LMM all score an AUROC of 0.47–0.57 across three benchmarks. A plain trainable MLP reaches 0.70–0.97, so model capacity is not the bottleneck; an input ablation shows that the residual signal comes from question-answer statistics, not image grounding. We then ask what survives. QCE-Net, an interpretable detector that pools image patches with both question and answer as attention queries, matches or exceeds the MLP baseline, recovering the grounding signal that frozen encoders can provide. LLaVA-LoRA, which fine-tunes the medical LMM with LoRA to emit per-sample grounding judgments, goes further: it reaches 0.885 AUROC on VQA-RAD, the first detector under a leakage-free protocol to clear the clinical bar of 0.8 that no frozen-CLIP method reaches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.