acceptodds
Under review as a conference paper at ICLR 2027

MedFrame-R1: Reward Gating & Informative Sampling for Medical Multimodal Reasoning

Abstract

Medical image interpretation requires combining visual evidence with clinical knowledge, making multimodal large language models (MLLMs) a promising approach to medical question answering and image understanding. To improve the accuracy of their answers, these models can be further trained using reinforcement learning (RL), which rewards desirable responses. However, effective training and reliable evaluation remain challenging due to limited learning signals, unreliable rewards, and evaluation bias. We introduce MedFrame R1, an open-source, reproducible framework that brings together data selection, reward design, and supervised initialization to study and improve medical multimodal RL. We also introduce an evaluation protocol that combines deterministic scoring with independent judgments from three openly available models, reducing reliance on any single judge. Across 17 medical benchmarks, our 4B model achieves a macro score of 64.8, outperforming our reproduction of MediX-R1 4B (62.9) by 1.9 percentage points. These contributions support systematic investigation of RL training and more reliable assessment of its benefits.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.