acceptodds
Under review as a conference paper at ICLR 2027

Noise-Robust Reinforcement Learning with Verifiable Rewards for Open-ended Medical Multimodal Reasoning

Abstract

Reinforcement learning with verifiable rewards (RLVR) has shown strong potential for improving the reasoning ability of multimodal large language models. However, most existing RLVR methods assume reliable labels for reward computation, making them vulnerable to annotation noise. This challenge is particularly prevalent in medical domains, where expert annotations are costly and often affected by diagnostic uncertainty, inter-observer variability, and complex clinical interpretation. In this work, we investigate noise-robust RLVR for open-ended medical multimodal reasoning and propose Noise-Robust RLVR(NR-RLVR), a framework that dynamically refines supervision signals during training. NR-RLVR consists of two key components. First, Reference-aware Pseudo-label Generation (RPG) improves pseudo-label reliability by aggregating predictions from multiple model rollouts and a reference-guided reasoning trajectory. Second, Confidence-gated Supervision Refinement (CSR) resolves conflicts between reference labels and pseudo-labels by selecting the more reliable supervision signal according to model confidence. Extensive experiments on text-only and vision-language medical benchmarks show that NR-RLVR consistently improves reasoning accuracy and robustness across different levels of label noise. Notably, under the challenging 90% noise setting, NR-RLVR improves the average performance of Qwen2.5-VL-3B by 7.47% over standard GRPO on vision-language medical benchmarks. Code is available in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.