acceptodds
Under review as a conference paper at ICLR 2027

IsoVGRBench: Video Reward Modeling Without Human Feedback

Abstract

Reinforcement learning post-trains video generators against reward models, so a generator improves only along the properties that its reward model perceives. Video reward models are trained and evaluated on human preferences between videos that often differ simultaneously on properties such as prompt alignment, visual fidelity, motion, and generator identity. As a result, aggregate accuracy does not reveal which video properties a reward model can reliably distinguish. We construct preference pairs by controlled interventions that change one property of a generated video while holding the rest fixed. This determines both the preferred side and the property under test without per-pair human labels, requiring human validation only once per intervention to confirm perceptibility. We propose ISOVGRBENCH, a benchmark of 8,000 such pairs across eight interventions, and show that current video reward models separate broad quality differences between generators but miss frame order. The best reward model reaches only 60.8% on reordered frames against 100% for humans, and rankings on existing benchmarks do not predict this failure. Intervention pairs can also serve as training data: our ISOVRM-Q2 outperforms VideoReward, which shares its base model but trains on human preferences, by 3.8 points on VBench-2.0 and provides better supervision for video generator post-training. Adding intervention pairs to VideoReward’s human-preference supervision further improves transfer. Because intervention pairs are automatically generated rather than collected, the benchmark and training data can be regenerated as generators evolve.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.