Sparse Evolution: Efficient Test-Time Self-Evolution for Few-Step Video Generation
Abstract
Few-step distillation substantially reduces the cost of video generation, yet test-time search remains expensive due to repeated full-video candidate generation. Unlike final video synthesis, candidate selection primarily requires reliable relative rankings, suggesting that generating every candidate in full may be unnecessary. We find that, after a single full-sequence warm-up step, sparse keyframe previews retain discriminative information and largely preserve the rankings of their full-video counterparts. Building on this observation, we introduce Sparse Evolution, an efficient test-time self-evolution framework for few-step video generation. Sparse Evolution adaptively selects keyframes using warm-up features and restricts subsequent denoising and decoding to the retained temporal positions, reducing computation. The resulting sparse previews guide candidate selection and provide evidence for iterative prompt refinement, with full-video generation completed only for the selected prompt–seed pair. Across three few-step video generators, sparse previews achieve mean per-prompt Spearman correlations of 0.75–0.81 with full candidates, while Sparse Evolution accelerates candidate generation by 2.00–2.25× over SVG2. On T2V-CompBench and VBench Semantic, Sparse Evolution achieves 1.74–1.91× end-to-end speedup while retaining 91% of the quality gains from test-time search.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.