acceptodds
Under review as a conference paper at ICLR 2027

OpenWan-Light: High-Quality Data Enhancing and Accelerating Video Generation

Abstract

Recent advances in video generation have greatly improved visual quality and temporal coherence, but the growing scale of modern models makes inference increasingly expensive, limiting their use in real-time applications. While models and acceleration algorithms are becoming increasingly open, the data used to train high-quality accelerated video models remain comparatively underexplored. We introduce a 450K real-video dataset centered on rich and diverse motion while preserving fine visual detail and strong text alignment, curating from 4.55M clean, standardized candidate clips. Our data-curation pipeline combines standard filtering and agentic VLM-based review to identify and balance dynamic video content. To demonstrate its effectiveness, we further build a complete acceleration pipeline on Wan2.2 A14B, combining supervised fine-tuning and few-step distillation with a generator-matched one-step refiner for high-resolution detail recovery. Without quantization or sparse attention, our system achieves acceleration over the reference generation pipeline and generates a 5-second 480p video in approximately 4.9 seconds on single NVIDIA H100, while maintaining strong generation quality and user preference. We will publicly release the dataset, model checkpoints, and complete training and inference code to support reproducible research on efficient and dynamic video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.