acceptodds
Under review as a conference paper at ICLR 2027

SAFER: When Text Beats Vision in Radiology VLM Hallucination Detection

Abstract

A text-only classifier with no image access reaches AUROC 0.918 on a benchmark designed to test visual hallucination detection in radiology vision-language models (VLMs) — far exceeding every evaluated VLM, including the best real model (MedGemma-27B, AUROC 0.700). This is the central finding of SAFER (Structured Assessment of Faithfulness and Error in Radiology), an 8,294-sample hallucination-detection benchmark for chest radiograph VLMs (OpenI Indiana Chest X-ray and MIMIC-CXR), evaluated across eight VLMs and five constructed baselines, with a nine-part failure analysis that interrogates the benchmark and its evaluation metric rather than presenting them as settled instruments. Through four rounds of controlled elimination, we trace this leakage: grammar and collocation artifacts explain negligible variance; the closed 50-word swap vocabulary explains just over half (≈55%, ∆AUROC 0.229 of a 0.418 above-chance gap, chance=0.500); report length, a 53-term vocabulary list, and cross-source formatting do not explain a material fraction of the remainder under the controls we tested. A residual AUROC of 0.65–0.70 is observed separately in both evaluated sources, as an open limitation of this specific lexical-substitution injection design and a signal that cross-model AUROC comparisons on benchmarks built this way cannot, without further analysis, be attributed to visual grounding. Beyond the leakage analysis, we show that a composite metric redesigned to punish degenerate classifiers still rewards a near-constant classifier as top-ranked (a candidate MCC-floor gate is demonstrated, though not independently validated); that four of eight evaluated VLMs collapse to an almost-exactly-constant output while one additional model (MedGemma) is the only non-degenerate VLM and shows the strongest evidence of image-dependent behaviour; and that calibration terms in composite metrics can promote a text-only baseline above all real VLMs under calibration-emphasised weighting. We release all raw metrics, per-sample outputs, and the nine-experiment analysis pipeline for independent verification, and recommend eliminative, control-paired validation as standard practice for benchmarks that use model self-report as a proxy for visual clinical reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.