Is LLaVA-Med Visually Grounded?
Abstract
Medical Vision-Language Models (Medical VLMs) offer a scalable pathway to support radiological workflows by automating report generation and clinical question answering. LLaVA-Med is one of the most widely adopted Medical VLMs. It is instruction-tuned on PubMed Central, enabling the original LLaVA model to adapt to the medical domain. However, we identify that LLaVA-Med is susceptible to a critical failure mode: rather than learning genuine visual-textual coherence, it often exploits biased language priors, effectively memorizing statistical patterns in clinical text and replaying them at inference time with limited regard to the actual image content. This undermines diagnostic reliability and poses a direct risk to patient safety in real-world clinical deployment. In this paper, we first propose a diagnostic framework to systematically identify and analyze the underlying causes of this failure mode. Building on these insights, we develop interventions that mitigate linguistic bias and encourage LLaVA-Med to generate responses that are more faithfully grounded in visual evidence. Results on in-distribution and out-of-distribution datasets like VQA-RAD, SLAKE and PathVQA support the claim.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.