acceptodds
Under review as a conference paper at ICLR 2027

Video-Rubrics: Fine-Grained Rewards for Faithful Embodied Video Generation

Abstract

Video diffusion models have made substantial progress in visual quality, yet instruction adherence remains challenging in embodied video generation. Semantic rewriters enrich the conditions supplied to video generators, but rewards based only on rewritten requirements can favor plans that omit difficult parts of the original request. We introduce Video-Rubrics, a reinforcement learning framework that jointly optimizes a prompt rewriter and a video generator through fine-grained visual criteria. We construct Video-Rubrics-60K, a training set of image-text requests spanning general scenes and robotic instructions, with atomic requirements annotated independently of candidate rewrites. Original-prompt rubrics supervise both requirement preservation and visual realization for the rewriter, while condition rubrics evaluate how faithfully the generator follows its actual input. A frozen vision-language verifier assesses each requirement using timestamped video evidence, with soft temporal windows that preserve action context. We normalize item-level rewards within each policy's comparison groups and use the resulting advantages to update both policies from the same sampled videos. Experimental results show that Video-Rubrics improves instruction adherence across three video backbones, with gains on multi-entity interactions and long-horizon tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.