acceptodds
Under review as a conference paper at ICLR 2027

Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

Abstract

Recent video diffusion models advance rapidly by scaling model size, data, and compute, but this trend leaves open a basic empirical question: at a fixed compute budget, which design choices across the video-diffusion stack actually pay for themselves? We address this question with Open-Sora 2.0, a commercial-level video generation model trained for only , 5 to10 times lower than comparable systems such as MovieGen and Step-Video-T2V, that we treat as a deliberately budget-bound empirical apparatus rather than a recipe report. Within this regime we report a quantitative answer to three coupled questions about cost-bounded video diffusion training: (P1) what makes a video latent space easy to denoise, where we find that higher autoencoder channel dimensions slow diffusion convergence despite improving reconstruction, and that DINOv2 latent distillation accelerates training on high-compression autoencoders; (P2) why image conditioning is the key efficiency lever, where image-to-video (I2V) models adapt to new autoencoders far more efficiently than text-to-video (T2V) models, and motion learned at low resolution transfers cleanly to high resolution through I2V fine-tuning; and (P3) how to combine image and text guidance at inference, where decoupling classifier-free guidance with a time- and frame-dependent image-guidance schedule resolves the otherwise conflicting requirements of motion fidelity and semantic alignment. Each claim is paired with a controlled experiment, a natural ablation, or a statistical test: a controlled FLUX-isolation experiment on VBench, a 22-run training-log analysis treated as natural ablations, and a human evaluation reporting Fleiss' and bootstrap 95% confidence intervals. Open-Sora 2.0 is comparable to leading systems including the open-source HunyuanVideo and the closed-source Runway Gen-3 Alpha. We fully open-source the model, training code, and data pipeline so every claim is reproducible.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.