acceptodds
Under review as a conference paper at ICLR 2027

A Reliability-Aware Evidence Organization Framework for Retrieval-Augmented Medical Visual Question Answering

Abstract

Medical large vision-language models (Med-LVLMs) have advanced medical visual question answering (VQA), but they can still produce answers unsupported by the image or clinical evidence. Retrieval-augmented generation (RAG) improves factual grounding by introducing radiology reports, yet visually similar examinations do not necessarily yield question-relevant or clinically consistent evidence. Report-level retrieval therefore suffers from two fundamental mismatches: retrieved reports may fail to cover the queried finding, and the retrieved content may contain irrelevant, uncertain, or conflicting statements. We propose \method, a novel factual grounding framework for medical VQA, which converts retrieved reports into question-specific clinical evidence. Without altering the image retriever, \method decomposes reports into structured evidence units encoding clinical findings, anatomical locations, and assertion status. It then uses a question-conditioned reliability estimator to jointly assess target relevance, uncertainty, and cross-evidence inconsistency. A risk-guided organizer subsequently selects and composes lower-risk, question-aligned evidence into a compact, clinically coherent context for the Med-LVLM. By separating evidence acquisition from evidence utilization, \method enables the Med-LVLM to reason over organized question-specific evidence rather than unfiltered reports. Its key contribution is not stronger retrieval, but a representation-level bridge that transforms retrieved reports into more reliable evidence for clinical reasoning. Under a unified evaluation protocol, \method reaches F1 scores of 88.60%, 88.89%, and 56.05% on IU-Xray, MIMIC-CXR, and PadChest-GR, with larger gains when retrieved reports already contain mixed but question-relevant evidence.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.