acceptodds
Under review as a conference paper at ICLR 2027

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification

Abstract

Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exploiting language priors over fine-grained visual verification. While rubric-based evaluation mitigates these biases in text-only settings, extending it to multimodal tasks is bottlenecked by the complexity of visual reasoning. The critical differences between responses often depend on instance-specific visual details. Robust evaluation requires dynamically synthesizing rubrics that isolate spatial and factual discrepancies. To address this, we introduce DeltaRubric, an approach that reformulates multimodal preference evaluation as a plan-and-execute process within a single MLLM. \dr operates in two steps: acting first as a , the model generates a neutral, instance-specific verification checklist. Transitioning into a , it executes these self-generated checks against the image and question to produce the final grounded judgment. We formulate \dr as a multi-role reinforcement learning problem, jointly optimizing planning and verification capabilities. Validated on Qwen3-VL 4B and 8B Instruct models, DeltaRubric achieves solid empirical gains, improving base-model overall accuracy on VL-RewardBench by (4B) and (8B) points and consistently surpassing no-rubric baselines. Beyond accuracy, DeltaRubric outperforms compute-matched inference scaling (verbose CoT and best-of-), preference learning (DPO), and automated prompt optimization while generating fewer tokens, generalizes to a different backbone InternVL3 and to scalar rewards, and as a frozen reward model improves downstream policy training. The results demonstrate that decomposing evaluation into structured, verifiable steps leads to more reliable and generalizable multimodal reward modeling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.