When Preference Margins Hide Text-Only Shortcuts: Visual-Conditional Margin Learning for Multimodal Reward Models
Abstract
Multimodal reward models (MM-RMs) should assess response quality using cross-modal information, yet often rely on text-only shortcuts, leading to underuse of visual evidence and potentially impairing reward judgments and downstream optimization. By analyzing text-only preference signals in pretrained representations and their evolution during reward-model training, we find that the standard Bradley–Terry (BT) objective widely used to train reward models from pairwise preferences, despite improving multimodal preference prediction, further amplifies pre-existing textual preference signals and leaves the contribution of visual evidence to the pairwise preference margin unsupervised. Based on these observations, we propose visual-conditional margin learning (VCM), a unified framework with two levels of visual-conditional supervision: visual-gain supervision from text-only to multimodal conditioning, and stronger evidence-flip (EF) supervision enforcing preference-margin changes under local reversals of decisive visual evidence while keeping the question and responses fixed, thereby cancelling the shared text-only component. We evaluate our method across multiple MM-RM backbones and multiple multimodal preference benchmarks, showing improved reward-model performance, reduced text-only shortcut dependence, and better downstream response selection in Best-of- evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.