InstaGen: Group-Shared Importance Sampling for Fast Video Generation
Abstract
Diffusion Transformers (DiTs) are the dominant architecture for video generation but suffer from significant latency due to the quadratic complexity of attention. Block-sparse attention reduces this cost by computing only the top-k attention blocks per query, or top-p, which varies k to meet a target attention-mass threshold, and skipping the rest. Achieving high generation quality under a fixed block budget requires solving two problems: (1) Top-k block selection: the blocks carrying the most attention mass must be identified accurately for each query cluster, but existing methods rank blocks by the score of a single cluster centroid, which misranks them whenever the queries within a cluster disagree. (2) Skipped block recovery: the blocks below the top-k boundary still carry attention mass, and existing methods either drop them, incurring information loss that grows as k shrinks, or compensate for them with per-block key/value centroids, whose reconstruction error is determined by the clustering and does not decrease as more compute is spent. In this paper, we propose InstaGen, which addresses both. For selection, InstaGen product-quantizes queries and keys so that block scores are cheap enough to compute for every query rather than only the centroid, and each query votes for its top-k blocks; the cluster attends to the blocks with the most votes. For recovery, InstaGen draws a small random sample of tokens from the blocks below the top-k boundary, shared across all queries in the cluster, and computes them together with the top-k blocks in a single attention kernel, with sampled scores reweighted to account for the full skipped region. The sample size is a direct budget: reconstruction error decreases as more tokens are sampled, at a cost that is a small fraction of the top-k computation. InstaGen is training-free and requires no auxiliary predictor. On Wan2.1-1.3B and Wan2.1-14B, InstaGen improves PSNR over SVG2 by 2.9 dB and 1.2 dB at matched attention density and by 1.8 dB and 0.9 dB at matched wall-clock time, and reaches SVG2's best quality 1.13× and 1.20× faster, respectively, while continuing to improve at densities where SVG2 has plateaued.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.