AdvFidelity: Unifying Preference Alignment and Few-Step Image Generation within Reinforcement Learning
Abstract
Practical image generation demands both alignment with human preference and few-step efficiency, yet the corresponding post-training procedures, reinforcement learning (RL) and few-step distillation, remain difficult to unify. We observe that few-step RL training carries an intrinsic distillation effect. Although few-step sampling on its own collapses into incoherent structures, generations evolve toward structural coherence as RL training proceeds. However, this effect could not remove grid-pattern artifacts, and preference rewards are inaccurate on such low-level degradations. Therefore, we introduce a fidelity reward that makes the intrinsic distillation effect explicit and reliable by supervising the removal of these artifacts. In AdvFidelity, this reward is realized as a discriminator that is updated dynamically along with the generator. The discriminator is trained on freshly constructed pairs at each training iteration. The negatives in each pair are drawn from the current generator and differ from their clean counterparts primarily in artifacts, so the discriminator keeps tracking the generator's evolving failure modes. The preference and fidelity rewards are differentiable and are optimized along the sampling trajectory used at inference via leap-based gradient propagation. In this way, a single RL stage performs preference alignment and few-step distillation simultaneously. AdvFidelity generates preference-aligned images in 8 steps without classifier-free guidance, matching or exceeding the quality of full-step CFG-guided generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.