MedVESS: From Seeing to Understanding Medical Images with Training-Free Latent Steering
Abstract
General Vision-Language Models (VLMs) offer a promising foundation for medical image understanding without requiring task-specific model design. However, reliable medical image understanding remains challenging, with insufficient reliance on medical visual evidence and limited medical semantic interpretation of the visual content. To address these limitations, we propose MedVESS, a training-free latent steering framework that constructs two complementary, modality-specific steering directions. Visual Evidence Steering (VES) contrasts original medical images with fully masked counterparts to strengthen reliance on medical visual evidence, while Clinical Semantic Steering (CSS) contrasts complete medical reports with semantic-depleted counterparts under the same visual input to reinforce associations between medical visual patterns and corresponding clinical concepts. The two steering vectors are constructed offline and jointly injected into a selected intermediate layer at inference time without updating model parameters. Experiments on four medical VQA benchmarks spanning radiology, fundus photography, and pathology show consistent improvements across two general VLMs, achieving competitive performance against several specialized medical VLMs without medical-domain fine-tuning. These results highlight the potential of training-free latent steering as a lightweight approach to improving reliable medical image understanding in general VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.