EVIDENCE THAT CAN CONTRADICT: WHAT REPORT- AUDITING METRICS ACTUALLY MEASURE IN MEDICAL VLMS
Abstract
As vision-language models (VLMs) begin drafting radiology reports in clinical settings, their safety relies on audit signals that flag which statements require physician review. The community typically evaluates these signals by pooling all generated claims into a single ranking and reporting the AUROC. We show that this standard metric is dominated by a prior that requires no visual information at all. Pooled claim-level AUROC decomposes exactly into within-concept and between-concept comparisons; on MIMIC-CXR, to of the pairs it scores compare different concepts, where the ordering is largely settled by knowing which findings are generally stated correctly. We derive a prior ceiling - the maximum score achievable by any predictor constant within a (concept, polarity) group, computable directly from claim counts without a model or an image. Across five generators from M to B parameters, this ceiling exceeds the strongest visual audit signals, including a recently published black-box auditor, by to AUROC. On the standard CheXbert-14 evaluation, the between-condition share is to and the ceiling is to , so reported pooled AUROCs should be read against this ceiling, not . Because the ceiling is an oracle, we also rank claims by training base rates alone; the visual audit signals beat this realisable baseline, but by a surprisingly narrow margin. We propose a stratified metric that removes this prior, under which the base-rate predictor scores exactly , and show that it reorders current auditing methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.