Diffusion-ART: Efficient, Stable and Scalable Reinforcement of Diffusion Model
Abstract
Reward-guided fine-tuning adapts diffusion models to preferences and task-specific objectives beyond pretraining. Yet scaling it remains difficult: reverse-process methods such as Flow-GRPO explore locally and have poor credit assignment, limiting their efficiency, while forward-process methods such as DiffusionNFT can be unstable without carefully tuned off-policy training. We show empirically and analytically that the instability comes from a hidden forward–reverse consistency force beyond the ideal NFT-style reward guidance, and that the restoration gradient is central to the stabilizing effect of off-policy updates. We introduce DiffusionART (Diffusion Reinforcement with Aggregation-First Reward Tilting), a simple yet effective method that systematically cancels the unstable self-consistency force and provides adaptive restoration. It requires minimal hyperparameter tuning, remains effective with fewer samples per prompt, and naturally supports pairwise reward models (DiffusionART-Pairwise). Experiments spanning text-to-image, text-to-video, and image-to-video generation show that DiffusionART is highly efficient and stable. For example, on HunyuanVideo-13B text-to-video generation, DiffusionART matches the plateau training rewards of DiffusionNFT and Flow-GRPO-Fast with 11.19× and 11.23× fewer GPU hours, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.