VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation
Abstract
Visual autoregressive (VAR) models offer fast, high-quality text-to-image generation but still fail at attribute binding and spatial relations. Existing VAR test-time methods search sampled trajectories or apply generic logit guidance but do not directly optimize compositional constraints; gradient-based diffusion methods do not transfer to stateful, discrete, multi-resolution VAR sampling. We introduce (sual Autoregressive emantic est-time lignment), the first gradient-based test-time framework for compositional alignment in next-scale autoregressive image generation. Built on Infinity, VISTA optimizes intermediate representations through the frozen transformer without parameter updates or training. It stabilizes optimization across scales and exposes a common interface into which any differentiable constraint on cross-attention can be plugged. Across two benchmarks and model scales, VISTA improves every targeted category, raising the mean by nearly 20% on a 2B backbone and almost 6% on an 8B backbone, especially on spatial relations. In human evaluation, VISTA wins 83% of decisive compositionality comparisons by plurality vote, with no significant preference in perceived quality; independently, ImageReward scores its outputs 20% higher. VISTA enables the 2B model to surpass a backbone four times its size, showing much of the compositional gap is recoverable at test time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.