acceptodds
Under review as a conference paper at ICLR 2027

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

Abstract

Extending the generation horizon of video diffusion models to long sequences remains a long-standing and important challenge. Existing training-free approaches either extend bidirectional models through architecture-specific modifications that degrade over long horizons, or rely on autoregressive generation that accumulates drift through exposure bias and produces repetitive motion. We instead formulate long video generation as an inverse problem: a long video is decomposed into a chain of short chunks, each constrained to be consistent with its neighbor's denoised estimate, which serves as the measurement. Following diffusion inverse solvers, we solve this problem with a manifold-constrained update on the denoised estimates, which we call Tweedie matching; it reduces to a closed-form per-frame solution and enforces temporal consistency while restricting the correction to the clean data manifold of the pretrained model. Because deterministic ODE trajectories tend to revert to their independent paths, we further propose stochastic early-phase sampling, which injects fresh noise after each correction in the high-noise phase to synchronize the windows, before switching to deterministic sampling to preserve fine-grained visual fidelity. The resulting framework is training-free and architecture-agnostic. Applied to various video generation models, it generates videos several times longer than the native window while outperforming both training-free and autoregressive baselines on VBench and long-range drift metrics, and extends without fine-tuning to audio-video joint generation and text-to-3DGS.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.