CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
Abstract
Diffusion-based video generative models achieve strong photorealism and motion fluency, yet fall short in professional production, where conditions are abstract, sparse and often mutually conflicting, e.g., storyboard sketches and clay-render clips that specify the creative intent rather than the pixels. We present **CogOmniControl**, a closed-loop reasoning-generation-verification framework for controllable video generation. The reasoner, CogDirector, is a specialized VLM trained via SFT and RFT on authentic production data; it cognizes the creative intent behind sparse and abstract conditions into a dense production plan, whose representations are injected into generator as semantic guidance through in-context conditioning. The same frozen CogDirector serves as the reward model during generator RFT, anchoring training and test-time verification to a consistent interpretation of creative intent. At test time, it further emits an evaluator harness in the same forward pass to select among multiple sampled candidates via VLM judges and specialized tools, i.e., test-time scaling. We theoretically characterize this verification process, deriving bounds on the misranking probability and expected selection regret of best-of-N selection, and showing that input-adaptive evaluation admits tighter guarantees when it reduces effective evaluation noise while preserving the target quality. We also introduce CogReasonBench and CogControlBench built from real professional workflow data with genuine creative intent. Experiments on the two benchmarks show that CogOmniControl surpasses existing open-source models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.