acceptodds
Under review as a conference paper at ICLR 2027

CamRM: Towards Reliable Camera-Motion Control in Video Generation with Motion-Grounded Rating Models

Abstract

Test-time scaling provides a practical way to improve camera-motion control in video generation: sample a clip, rate its alignment with the request, and revise the prompt if needed. This process requires a reliable rating model (RM) to determine when to stop. Geometry-based methods can struggle when object motion dominates the view, while fine-tuned visual language models (VLMs) often fail to adapt to synthetic videos. Our benchmark of human-labelled real and synthetic videos shows that these methods struggle to judge camera-motion alignment. Strong performance on real video does not ensure reliable judgment on generated clips: a model tuned only on real footage reaches % balanced accuracy on real video but drops to % on Omni clips. We introduce CamRM, which combines video frames with motion cues to judge camera-motion alignment. We develop CamRM on both the proprietary Gemini model and the open-source Qwen model, with consistent improvements of over % across both. We also develop a training-free agentic model (CamRM-A), showing that our approach requires no fine-tuning. Our CamRM achieves % accuracy and improves balanced accuracy over the strongest baseline by %. In the motion-guided refinement loop, using CamRM as RM improves human-judged alignment by % for Veo 3.1 and % for Omni Flash at the generation cost, reaching % alignment in the Omni Flash T2V case. At generations, the loop reaches % pooled candidate-set oracle alignment versus % for i.i.d. resampling, and matches its -generation oracle rate with three generations, without generator training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.