acceptodds
Under review as a conference paper at ICLR 2027

Exploring Unsupervised Multimodal Reward Modeling

Abstract

Vision-language reward models are trained on annotated preference pairs and ranked by benchmark accuracy; the pairs are costly and noisy to curate from human annotations, and the number is trusted more than it is understood. We present CRB, a reward model that uses no human preference label: a linear head read out of an intermediate layer of a frozen Qwen3-VL backbone, trained to prefer a web caption over a same-image corruption of it. The readout layer is the decisive choice, worth up to 5.8 points over the last layer. With it, one recipe transferred without re-tuning matches GPT-4o at 2B and 8B on all three benchmarks we test, by 0.5, 6.6 and 2.4 points on Multimodal RewardBench, VL-RewardBench and MMRB2 reasoning, from a few thousand trained parameters at about 0.1% of fine-tuning compute. At an equal pair budget, corruptions teach a frozen head 3 to 5 points more than curated RLAIF-V pairs. We then show what the number does not certify. The objective is convex over frozen features, so no training-side change moves its plateau. In Best-of-N the probe beats majority vote by 7 to 11 points only where the generator is confidently wrong. Two probes with the same score, trained on halves of the data split by whether CLIP can see the corruption, improve by 6.4 points on hallucination judging. What did improve deployment was correctness labels minted by cheap verifiers from the policy's own candidates: one bootstrap turn reduces vision hallucination by 5.1 CHAIR points and lifts image-generation pass rate by 5.4. With the combination of corruption, read-out and bootstraping, out method makes unsupervised MM reward learning cheaper and more efficient.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.