Gradient-Space Evolution Strategies for Diffusion Model Alignment
Abstract
Reinforcement learning aligns diffusion models with learned reward functions, but repeated trajectory evaluation and gradient computation make training expensive. We propose Gradient-Space Evolution Strategies (GS-ES), which organizes alignment around population-based exploration at individual denoising steps. GS-ES perturbs a shared transition mean, evaluates candidate states through gradient-free ODE continuations, and aggregates reward-weighted score-function updates. Selected candidates guide subsequent exploration, while a timestep-dependent budget allocates evaluations according to the remaining rollout length. For Gaussian transitions, we derive an exact identity expressing the local update as a Jacobian pullback of weighted action perturbations. This connects evolution-inspired exploration to gradient-based parameter updates without a reference-policy KL penalty. Experiments with FLUX.1-dev and Wan2.1-T2V-1.3B show higher scores than the evaluated baselines across the listed image and video metrics. On FLUX.1-dev, GS-ES achieves a PickScore of 24.27, compared with 23.33 for DanceGRPO, and reaches the baseline's peak reward in approximately one fifth as many training steps.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.