acceptodds
Under review as a conference paper at ICLR 2027

High-Quality Demonstrations in Diffusion Reinforcement Learning

Abstract

Reinforcement learning (RL) improves diffusion models by rewarding generations that satisfy task objectives. Yet task rewards capture only part of visual quality, while the model's generations may continue to fall short of task requirements. High-quality external prompt–image pairs, which we call demonstrations, provide holistic visual supervision beyond what task rewards capture. To incorporate this supervision into RL, we introduce Demonstration-Augmented NFT (Demo-NFT), which combines two sources of image targets in DiffusionNFT's objective: prompt-matched external demonstrations and native samples generated by the diffusion model being optimized. Their rewards jointly determine the advantages used to weight these targets. However, large reward gaps can skew this weighting toward demonstrations. Native-referenced reward capping calibrates demonstration rewards against those of native samples. With 5% demonstration targets, Demo-NFT improves QJudge, an auxiliary metric of image quality, by 8.84 points on average over NFT across FLUX and SD3.5 on compositional generation and text rendering, while achieving better or comparable task performance. Using only the task reward, Demo-NFT also outperforms a multi-reward NFT baseline combining task and HPSv3 rewards on the task metric, QJudge, and PickScore in all four settings. We further examine supervised mixing and shared SFT initialization, and extend the evaluation to Qwen-Image and SenseNova. These results show that Demo-NFT improves multiple aspects of image quality using a single task reward.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.