acceptodds
Under review as a conference paper at ICLR 2027

If Medical Concepts Are Still There, Why Do Med-LVLMs Get Them Wrong?

Abstract

Why do medical large vision–language models (Med-LVLMs) get medical concepts wrong when those concepts remain recoverable from their internal representations? We investigate this gap between concept availability and answer correctness by tracing medical concepts through the visual encoder, vision–language connector, and language decoder. Linear probes reveal no sustained loss of concept decodability, while representational analyses indicate that encoder features are preserved across the connector. The decoder, however, exhibits differences in how available information is used: incorrect answers show weaker attention to target regions, and correct and incorrect answers become increasingly separable in middle-layer MLP activations. Activation patching further identifies middle-to-late MLP components whose clean outputs restore support for concept presence after target-region masking. These findings suggest that errors can arise from a failure to use available medical information, rather than from its loss alone. Guided by this evidence, we introduce MedSteer, which uses concept-specific directions derived from original and target-masked images to edit decoder activations, with a presence probe gating each edit. Across three Med-LVLMs, MedSteer improves average VQA accuracy by – percentage points over the unedited models. Our results identify a gap between representing medical concepts and using them to answer correctly, and demonstrate that concept-guided decoder editing can help bridge this gap.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.