acceptodds
Under review as a conference paper at ICLR 2027

VIGOR: Enhancing Reliability and Reasoning in Medical Report Generation through Reinforcement Learning

Abstract

Multimodal Large Language Models have shown strong performance in vision–language tasks, yet their application to medical report generation remains limited by insufficient visual grounding, opaque reasoning, and reliance on supervised fine-tuning, which primarily learns to imitate reference reports rather than explicitly optimizing clinically grounded objective. We propose VIGOR, a two-stage reinforcement learning framework for radiology report generation. In the first stage, we introduce a verifiable IoU-based reward and optimize the model using GRPO, enabling the trained model to explicitly align textual descriptions with region-level visual evidence, thereby enforcing accurate localization of abnormalities and improving interpretability. In the second stage, we extend GRPO to report generation by incorporating disease label and LLM-as-a-Judge reward to promote diagnostic accuracy and coherent reasoning. This two-stage optimization enables the generation of reports that are visually faithful, diagnostically consistent, and grounded in interpretable reasoning. Extensive experiments on the resulting model, VIGOR-4B, demonstrate reduced hallucinations, improved disease identification, and stronger alignment with verifiable evidence. Our results highlight the importance of explicit visual grounding and reward-guided optimization for developing reliable and faithful medical report generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.