Delta-Vision Reward: A Distribution-Aware Multimodal Reward Model
Abstract
Reward models are central to aligning multimodal large language models, yet most existing methods reduce each response to a single scalar score. This simplification ignores disagreement in evaluation and weakens supervision on ambiguous or contested samples. We present Delta-Vision Reward (DVR), a distribution-aware multimodal reward model that predicts a Gaussian reward distribution together with a structured rationale. The predicted mean captures the central preference signal, while the variance models disagreement induced by a teacher reward ensemble. DVR is trained in two stages. We first perform supervised distributional imitation, teaching the model to match teacher-derived reward statistics and explanatory reports. We then refine pairwise preference behavior by focusing only on response pairs where the student disagrees with the teacher consensus or exhibits insufficient preference margin. This targeted refinement improves learning efficiency by concentrating supervision on hard cases. Experiments on VL-Reward-Bench and MM-RLHF-Reward Bench show that DVR is a strong open-source reward model, with clear gains in preference prediction, confidence-sensitive evaluation, and distribution alignment. We further show that DVR produces actionable feedback that helps improve responses from a base multimodal model. These results suggest that reward modeling for multimodal alignment benefits from preserving uncertainty rather than collapsing it into a scalar target.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.