acceptodds
Under review as a conference paper at ICLR 2027

Phy-Judger: Learned Physics Rewards for Joint Prompt and Video RL

Abstract

Video generators produce visually convincing clips that still violate basic physics, such as a hand passing through a tabletop. Prior reinforcement learning (RL) methods for physical video generation reward the generator with hand-designed verifiers, each checking one property from object tracks, 3D estimates, or simulation states. These measurements are unreliable in open-world videos, and violations of other properties go unpenalized. We instead learn the reward from a physics preference dataset of 10k prompt–video pairs. The prompts cover 25 sub-scenes in five categories: human motion, indoor and everyday objects, industry and robotics, traffic and vehicles, and nature and materials. The label of each pair is known from how the pair is built, e.g., a real clip over a generated one, so annotators only write rationales. A rationale scores each video on six physical dimensions, such as mechanics and geometry, and lists each violation with its time interval and bounding box. On these data we train Phy-Judger, a VLM reward model that writes a rationale before scoring a video, so each score is tied to specific violations. We also train a prompt policy jointly with the video policy; it rewrites each prompt to add missing physical conditions, such as material, contact, and support. Joint RL with Phy-Judger raises the VBench-2.0 Physics score of MiniMax-H3 from 75.80 to 78.83 and its PhyWorldBench physical commonsense score from 0.155 to 0.203. We will release the code, checkpoints and datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.