Contrastive Decoding Is Classifier-Free Guidance: Unifying Visual Guidance for Vision-Language Models
Abstract
Classifier-free guidance is standard in image and video generation, yet is not commonly used in vision-language models. We show that the field has effectively used it without recognizing it: existing contrastive decoding methods are classifier-free guidance with different references, distinguished by the bias each reference carries. This unified formulation exposes a common design space for autoregressive generation. Exploring this space, we find that corrupted-image references increasingly align with simply deleting the image, while weaker models and degraded attention provide no better guidance. Guidance operators from diffusion likewise reduce to vanilla guidance on different scales. The key distinction from diffusion lies in the prefix: an autoregressive reference must condition on some answer prefix, but it is unclear which. Existing methods use the model’s own prefix, leaking image information into the image-free reference, and rapidly collapsing the guidance signal. A prefix-free reference avoids this leakage but becomes stale and overguides long answers. Switching between the two balances these complementary failure modes. Empirically, our approach improves Qwen2.5-VL-7B by 5.7 points and Qwen3-VL-8B by 1.7 points across ten general benchmarks, outperforming four published contrastive decoders at 1.04× the cost of unguided decoding, compared with up to 2× for prior methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.