acceptodds
Under review as a conference paper at ICLR 2027

Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning

Abstract

Reinforcement learning (RL) post-training has substantially improved the reasoning accuracy of vision-language models, yet the resulting policies remain poorly calibrated. Terminal correctness rewards provide no gradient that penalizes confident errors more than uncertain ones, and no signal that ties confidence to the quality of the supporting visual evidence. This mismatch becomes especially severe under corrupted or ambiguous visual inputs, where models continue to report high confidence on wrong answers. We introduce **Ranking-Aware Calibration (RAC)**, a training-time framework that supervises confidence using two comparison signals that group-based RL already produces at no additional labeling cost. The *ranking-aware group loss* enforces that a better rollout receives higher confidence than a worse one within the same prompt. The *clean–corrupted pairwise loss* enforces that confidence attenuates as visual evidence degrades. The ranking signal requires the policy to distinguish between correct and incorrect rollouts and improves task accuracy on most evaluated backbones. Both losses require no external confidence annotations and integrate naturally with group-based RL post-training. We instantiate RAC on Qwen2.5-VL, Qwen3.5, and InternVL-3.5 backbones and evaluate on six multimodal reasoning benchmarks under clean and corrupted inputs. Empirical results show that the ranking-aware loss improves task accuracy on most evaluated backbones, while the pairwise corruption loss reduces calibration error. Their combination improves ECE and Brier score over Vanilla-RL on all four backbones and achieves best or joint-best calibration among the compared methods across all four, while improving accuracy on three of the four backbones. Our code is available at <https://anonymous.4open.science/r/RAC>.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.