acceptodds
Under review as a conference paper at ICLR 2027

QGR: Query Grounding and Refinement for Medical Visual Reasoning

Abstract

Visual grounding links textual expressions to relevant image regions, while grounded reasoning requires model predictions and explanations to be supported by visual evidence. This remains challenging for medical vision-language models (VLMs), which can produce plausible explanations that are poorly aligned with disease relevant regions, leading to hallucinated reasoning and unreliable interpre- tations. Existing medical visual question answering (VQA) approaches mainly rely on supervised fine-tuning (SFT) or reinforcement learning (RL) to improve answer generation and reasoning, but provide limited direct supervision for question aware spatial grounding. As a result, spatial representations are often learned indirectly through the language-modeling objective, making it difficult to align reasoning with relevant visual evidence. We propose QGR (Query Grounding and Refinement), a multimodal grounding framework that connects language reasoning with object level visual evidence. QGR has two key components: Question Conditioned Query Refinement, which adapts object queries to the question context to capture relevant evidence, and Query Level Grounding Supervision, which directly guides their semantic assignment and localization quality. Together, they establish question aware semantic spatial correspondence at the query level. QGR jointly learns an- swer generation, reasoning, semantic region prediction, bounding-box localization, and query level grounding. Experiments show consistent improvements over the corresponding SFT baselines. On Med4VQA, QGR improves Closed accuracy by +3.20–17.64 points, Box localization by +7.64–52.15 points, and Final scores by +4.18–10.11 points across four imaging modalities. The gains also transfer out of domain, reaching +14.49 Final points on SLAKE MRI and up to +35.00 Box points on GEMeX. These results highlight the benefit of query level supervision for spatially grounded medical reasoning. Full training details and anonymized code and dataset are provided for reproducibility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.