acceptodds
Under review as a conference paper at ICLR 2027

Confidence Inflation from RL Post-Training in Video Language Models

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves answer correctness in video multimodal LLMs, but what it does to confidence is unmeasured. We compare two checkpoints item by item based on the model’s option-letter distribution, which serves as the confidence readout, and limit the analysis to items that both answer incorrectly so that confidence changes are separated from changes in correctness. On that set, RL post-training systematically increases confidence on unchanged wrong answers, and we call this excess the calibration tax. The magnitude of the calibration tax varies across question types in a pattern that is associated with the post-training recipe and is not explained by item difficulty, and this effect replicates across three benchmarks and three model scales. At the item level, the tax appears as a scalar change in the readout that leaves the predicted option unchanged but can substantially alter the numerical confidence assigned to it. At a fixed confidence cutoff of 0.8, items whose predictions did not change are accepted for automatic handling 59% more often, and the error rate among them rises from 27.6% to 36.2%. We then ask where the tax can be controlled. Reward-side interventions can move the readout, but their effect is unstable across training runs. Because the tax also varies from item to item, a single global post-hoc correction removes only part of it, which motivates controlling confidence directly during RL. Our method, Reference Margin Anchoring (RMA), adds a two-sided penalty during RL that anchors the readout margin to the frozen reference already used by the KL term. In a 1200-step run, it removes 95% of the RL-induced calibration tax, at a cost of 0.4 accuracy points that fall within run-to-run variation, and returns the threshold curves close to their pre-RL level. Our results show that RL post-training can introduce substantial confidence inflation beyond its effect on answer correctness and that this increment can be controlled directly at the readout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.