VPG: Visual Prefix Guidance for Next-Scale Autoregressive Image and Video Generation
Abstract
Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time, creating exposure bias: early prediction errors can become part of the conditioning context and influence subsequent predictions. Existing remedies mainly modify the training or fine-tuning procedure to expose the model to perturbed or model-generated prefixes, while sampling-time guidance primarily strengthens external semantic conditions such as class labels or text prompts. We propose Visual Prefix Guidance (VPG), a training-free inference-time guidance method for visual autoregressive generation. VPG contrasts predictions under the generated and corrupted prefixes and favors candidates that are both sensitive to and positively supported by the generated prefix, encouraging continuations more consistent with the realized visual history. Across class-conditional image generation with VAR, text-to-image generation with Infinity, and text-to-video generation with InfinityStar, VPG improves generation quality without retraining or fine-tuning, reducing FID on VAR by 0.40 on average and improving benchmark performance on both image and video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.