GuidanceNFT: NFT Based Reward Fine-Tuning for Guided Diffusion Models
Abstract
Guidance and reward fine-tuning improve diffusion generation in complementary ways. Combining the two, however, can undermine both benefits: applying classifier-free guidance after DiffusionNFT (NFT) can reduce the optimized reward, while incorporating guidance during NFT fine-tuning may cause collapse. We identify a mismatch in NFT's regression center as a source of this instability. The center must match the conditional mean velocity obtained by re-noising generated samples; using a mismatched center introduces a bias that can rotate or reverse the reward-driven update. We propose GuidanceNFT, which learns this center online by flow matching on guided samples and corrects both NFT regression terms while preserving the policy's reward-driven displacement. The method supports explicit and intrinsic guidance without additional inference passes. On SD3.5-M, GuidanceNFT improves the optimized reward over standard NFT across three training objectives and outperforms NFT on all seven evaluation metrics for GenEval and OCR specialization. In particular, OCR reward increases from 0.971 to 0.975 while held-out GenEval success rises from 0.035 to 0.488, indicating improved prompt fidelity beyond the training objective. Across ImageNet generation, OpenWebText, and GSM8K reasoning, GuidanceNFT improves generation quality and task performance under intrinsic guidance. These results show that online center correction enables effective reward fine-tuning of guided diffusion models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.