acceptodds
Under review as a conference paper at ICLR 2027

MODALENS: Measuring Image Sensitivity in Report-Conditioned Medical VLMs

Abstract

A radiology report can already answer a clinical question, so it is hard to tellwhether a vision-language model also uses the image. MODALENS, a pairedimage-swap audit, measures how report availability changes image sensitivity:MedGemma-27B on 3,199 paired MIMIC-CXR cases from 293 patients, the 13finding-specific questions per case, each image replaced by one from another study,usually of the same patient, with question and report fixed. Under an explicit answerinstruction, the model’s generated answer changes on 4.43% of trials with the reportand 19.93% without it, a paired increase of 15.5 points (patient-clustered 95%CI [14.5, 16.5]), so report availability reduces image-swap sensitivity under thisprotocol, and substitutions also move continuous answer scores where the binaryprediction does not change. The direction replicates in two further general-domainlineages and on a second institution’s chest X-ray collection; two medical modelsread the substituted image at chance with or without the report. A cumulativeattention block on the report tokens restores image sensitivity when it begins at orbelow layer 22 and not from layer 28 upward, a causal depth bound rather than alocus. The labels are derived from reports, which limits conclusions about visualcorrectness. Code, exact prompts and a committed run record behind every numberare provided as anonymized supplementary material and will be released

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.