Free Diversity for Efficient Test-Time Scaling in Visual Autoregressive Models
Abstract
Visual Autoregressive (VAR) models generate images by next-scale prediction, and their earliest scales strongly shape the final image. This sensitivity has been taken to imply that these scales must remain text-conditioned. We revisit this assumption and find that text conditioning has little effect at the coarsest scales: sampling them unconditionally leaves mean quality nearly unchanged while substantially diversifying the generated images. Existing diversity-boosting methods for VAR improve best-of- performance at the cost of either latency or mean quality; in contrast, unconditioning the earliest scales incurs no extra latency and a smaller drop in mean quality while yielding higher diversity and best-of- performance. Building on this finding, we propose early attention pruning (\eap), an efficient verification method that prunes failing candidates at early scales using cross-attention signals, without decoding intermediate images. Our training-free pipeline outperforms existing diversity-boosting and test-time scaling methods on T2I-CompBench++ and GenEval2 with much less computation, and generalizes across VAR architectures and model sizes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.