acceptodds
Under review as a conference paper at ICLR 2027

Verifiable ≠ Valid: What Programmatic Physics Rewards for Video Generation Actually Optimize

Abstract

Post-training video generators with verifiable physics rewards, meaning programmatic quantities read deterministically off frozen perception models, has become a leading alternative to VLM judges. Such rewards are validated by showing that they separate real video from rule-corrupted video. We find that this validation systematically overstates their validity, that what the two rewards we could measure actually optimize is motion magnitude, and that removing the motion confound is necessary but far from sufficient. Across two human-rated datasets, 13 generators and two architectures, a reimplemented kinematic-residual reward correlates with peak optical flow at –. Once motion is partialled out it retains no association with human ratings in 0 of 12 generators under a single motion control, and a pure motion statistic containing no physics predicts human ratings better than it does. Optimization reproduces this causally: in two GRPO runs identical but for the reward, the motion-correlated reward drives peak flow down 46% (95% CI ) while a motion-orthogonal reward drives it up 67%, the arms separating monotonically in dose at , with three motion-orthogonal signals pushed in no consistent direction; a lower-fidelity second prompt set replicates it, with peak flow down 88% at the highest dose and the arms separating at all five doses. The identical test applied to a VLM physics judge answers oppositely: the judge retains 88.0% of its agreement with human ratings once motion is partialled out, against % for the programmatic reward. The dividing line is therefore not verifiability but whether the quantity is a motion statistic. Escaping the confound is not sufficient either: the most structured physics reward published to date, which grounds recovered 3D bodies in a physics simulator, does escape it ( with motion, against ) yet, tested outside its design domain, carries no measurable validity beyond motion on this rating (partial , ). We stop short of the last step. We do not show that the invalid reward makes videos worse for people: we have no human ratings on our own generations, the two runs disagree about the automatic judge, and treating a rising automatic physics score as evidence about human judgement is the very inference this paper is written against.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.