Commit-One: Decoupling Generation and Commitment for Interactive Video Generation
Abstract
Interactive video generation requires new conditions to influence ongoing playback promptly. A natural route is fine-grained autoregressive generation, but reducing chunk size can restrict joint temporal modeling and parallel computation. *Must fine-grained interaction therefore come at the cost of a large chunk size?* We introduce Commit-One, an algorithm-system co-design built on a different principle: **low-latency interaction requires fine-grained output commitment, not necessarily fine-grained generation.** By decoupling chunk size from output commitment, Commit-One preserves large-chunk generation while allowing future predictions to remain revocable. This enables chunk-wise autoregressive models to deliver fine-grained video streams without reducing their chunk size. The optimization target therefore shifts from making chunks smaller to making joint denoising faster: chunk size becomes a design choice balancing modeling capacity and computational efficiency, rather than a hard constraint on interaction granularity. The same principle opens the possibility of using even bidirectional teacher models directly for interactive streaming. Commit-One thus offers an alternative route to interactive video generation, centered on accelerating high-quality joint generation rather than shrinking chunk size.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.