CANVAS: PLAN THE SHOOT, NOT JUST THE STORY, FOR LONG VIDEO
Abstract
Video generators synthesize striking short clips, but assembling them into coherent multi-shot films remains difficult. Existing approaches plan the story and shots, then render each shot with the same conditioning policy and compute budget. We argue that coherence hinges on two shot-level decisions: what visual context to inherit and how much computation to spend. We introduce CANVAS, a training-free framework that makes both decisions before generation. It compiles a story into an executable shooting plan: pixels carry over only within continuous action, while structured state carries identities, objects, and narrative history across cuts. A film level budget assigns computation according to shot difficulty, and a verifier directs revisions to specific failures. On ViMax-Bench, with the language model, image generator, and video generator held fixed, CANVAS raises global consistency from 0.528 for ViMax and 0.452 for MovieAgent to 0.694, with VBench aesthetic and imaging quality on par with ViMax. It spends 0.76 times the denoising of a uniform 25-step pass on the same 35 stories; generation takes 136 minutes per film, versus 129 for ViMax and MovieAgent. The consistency gains are supported by three evaluator families and a human study, and persist without frame chaining. We will release the code to the public.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.