More than a Single Guidance Strength: Rethinking Classifier-Free Guidance in Next-Token Autoregressive Image Generation
Abstract
Classifier-free guidance (CFG) in next-token autoregressive image generation typically uses the difference between conditional and unconditional predictions as a fixed guidance direction, controlled by a single scalar strength. This raises a basic question: Is a single guidance strength sufficient to characterize how CFG should be applied? First, we view CFG as increasing conditional evidence while limiting its KL deviation from the model’s conditional prediction, and further study this trade-off under the fisher geometry. In this geometry, we decompose the CFG direction into two components with distinct effects on token prediction. One preferentially amplifies class-specific tokens, while the other preserves tokens favored by both conditional and unconditional predictions. We then investigate how the guidance direction should be combined for a given guidance strength and observe a systematic shift in the preferred balance: weak guidance favors more class-specific amplification, while stronger guidance requires greater preservation of jointly supported tokens. We further study how guidance should be applied for a given strength and composition, and find that the update strategy also matters: successive local updates become more useful with stronger guidance and greater emphasis on class-specific amplification. Together, these findings show that a single guidance strength is insufficient to characterize CFG, as its behavior also depends on how the guidance direction is composed and how the guidance is applied. Experiments across multiple autoregressive backbones validate these findings and show consistent improvements in both class-to-image and text-to-image generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.