acceptodds
Under review as a conference paper at ICLR 2027

EVIDENCE THAT CAN CONTRADICT: WHAT REPORT- AUDITING METRICS ACTUALLY MEASURE IN MEDICAL VLMS

Abstract

As vision-language models (VLMs) begin drafting radiology reports in clinical settings, their safety relies on audit signals that flag which statements require physician review. The community typically evaluates these signals by pooling all generated claims into a single ranking and reporting the AUROC. We show that this standard metric is dominated by a prior that requires no visual information at all. Pooled claim-level AUROC decomposes exactly into within-concept and between-concept comparisons; on MIMIC-CXR, to of the pairs it scores compare different concepts, where the ordering is largely settled by knowing which findings are generally stated correctly. We derive a prior ceiling - the maximum score achievable by any predictor constant within a (concept, polarity) group, computable directly from claim counts without a model or an image. Across five generators from M to B parameters, this ceiling exceeds the strongest visual audit signals, including a recently published black-box auditor, by to AUROC. On the standard CheXbert-14 evaluation, the between-condition share is to and the ceiling is to , so reported pooled AUROCs should be read against this ceiling, not . Because the ceiling is an oracle, we also rank claims by training base rates alone; the visual audit signals beat this realisable baseline, but by a surprisingly narrow margin. We propose a stratified metric that removes this prior, under which the base-rate predictor scores exactly , and show that it reorders current auditing methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.