Rethinking Classifier-free Guidance in Reward-based Diffusion Post-training
Abstract
Reward-based diffusion post-training updates model parameters from task rewards, yet the role of classifier-free guidance (CFG) in this process remains unclear. We identify an Inference–Learning Mismatch: guidance can improve a fixed model's outputs while stronger guided collection weakens subsequent learning. This contrast motivates using CFG to organize reward measurements rather than prescribe a training direction. We construct local reward-query coordinates from the conditional–unconditional residual and compare symmetric same-state interventions to identify reward-preferred changes. The resulting reward differences provide directional credit tied to a known intervention, complementing the sample-level credit of individual images. To combine these signals in a parameter update, we introduce Fiber-Augmented Negative-Aware Fine-Tuning (FANFT), which pulls local responses into parameter space and fuses them with global NFT credit according to agreement between disjoint estimates. Frozen-policy tests show that these coordinates predict held-out preferences better than matched random coordinates, while training comparisons connect directional writing to improved learning on condition-sensitive rewards. Across tasks, the gains depend on how well the coordinate represents reward-improving directions. These findings place reward–coordinate alignment at the center of CFG's post-training role: the residual organizes local queries, and global reward supervision complements the directional information they provide.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.