acceptodds
Under review as a conference paper at ICLR 2027

Delta-Vision Reward: A Distribution-Aware Multimodal Reward Model

Abstract

Reward models are central to aligning multimodal large language models, yet most existing methods reduce each response to a single scalar score. This simplification ignores disagreement in evaluation and weakens supervision on ambiguous or contested samples. We present Delta-Vision Reward (DVR), a distribution-aware multimodal reward model that predicts a Gaussian reward distribution together with a structured rationale. The predicted mean captures the central preference signal, while the variance models disagreement induced by a teacher reward ensemble. DVR is trained in two stages. We first perform supervised distributional imitation, teaching the model to match teacher-derived reward statistics and explanatory reports. We then refine pairwise preference behavior by focusing only on response pairs where the student disagrees with the teacher consensus or exhibits insufficient preference margin. This targeted refinement improves learning efficiency by concentrating supervision on hard cases. Experiments on VL-Reward-Bench and MM-RLHF-Reward Bench show that DVR is a strong open-source reward model, with clear gains in preference prediction, confidence-sensitive evaluation, and distribution alignment. We further show that DVR produces actionable feedback that helps improve responses from a base multimodal model. These results suggest that reward modeling for multimodal alignment benefits from preserving uncertainty rather than collapsing it into a scalar target.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.