acceptodds
Under review as a conference paper at ICLR 2027

Steerable but Not Readable: Fixation Heads in Medical Vision-Language Models

Abstract

Medical vision-language models (VLMs) write fluent findings but do not show which image region supports each sentence, so a clinician cannot check a claim against the scan. The usual remedy reads an attention or attribution map over the scan as a record of what the model used. Generalist VLMs contain a few attention heads that follow the region being described, yet it is unknown whether medical fine-tuning keeps these heads and whether their attention can serve as provenance. We test these fixation heads in medical VLMs in both directions, reading their attention as evidence and writing a region into it. Our evaluation across three model families and six corpora shows that (i) every model has fixation heads and supervised medical tunes inherit them from their base, (ii) the heads are steerable, as aiming them at a region changes what the model says, and (iii) they are not readable, as their attention barely follows the question, leaves the organ as the model answers and locates findings no better than random boxes, as do standard attribution maps. Anatomy-conditioned generation turns the write into provenance without training, writing one anatomical region into the heads per sentence so that the sentence depends on it, and adds clinical findings to chest reports. MedGaze, a low-rank adapter on the heads' query rows, learns the write, moving the heads onto the anatomy with no region at test and answering organ questions more accurately than training-free decoders. Our findings emphasize the need to write provenance during decoding rather than read it from attention maps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.