acceptodds
Under review as a conference paper at ICLR 2027

The Retrieval Paradox: Why RAG Fails Medical Vision-Language Models Across Architectures

Abstract

Retrieval-Augmented Generation (RAG) has emerged as a widely adopted approach for adapting general-purpose vision-language models (VLMs) to specialized medical domains at inference time. However, its effectiveness and safety for medical visual question answering (VQA) remain critically understudied. We present a comprehensive empirical evaluation of five RAG configurations across three VLM architectures and three medical VQA benchmarks (VQA-RAD, SLAKE, PathVQA), using a retrieval bank of 50,000 medical image-QA pairs. We uncover the Retrieval Paradox: RAG systematically degrades performance across all tested architectures and benchmarks without exception, dropping LLaVA-1.5-7B from 45.0% to 35.7% on VQA-RAD and LLaVA-1.6 from 45.68% to 36.81%. Critically, BLIP-2, which achieves 96.0% at , suffers the largest degradation: from 96.0% to 55.65% on VQA-RAD (). Rather than preventing failure, Q-Former bottleneck alignment creates structural inflexibility that amplifies context collapse when retrieved text is injected at the decoder. A controlled Oracle experiment recovers accuracy from 24.0% to 96.0% with perfect context, and a random text ablation ( vs. for medical RAG) isolates semantic content as the driver of failure rather than token length. Population-level attention analysis across the full test set (, paired -test ) confirms that retrieved medical text displaces visual attention by 28.0% relative. A modern BGE cross-encoder reranker provides no meaningful improvement ( vs. flat FAISS), establishing the failure as architectural rather than retrieval-quality-dependent. None of the tested decoder-based VLMs under standard prompt-prefixing is RAG-immune, with direct implications for the safe deployment of multimodal AI in clinical settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.