Diffusion Advantage Matching: Unified Diffusion Policy Optimization
Abstract
Diffusion policies can represent multimodal behaviors, but improving them online is challenging because their action likelihoods are intractable. Existing diffusion-RL methods address this difficulty through action gradients, likelihood approximations, or backpropagation through the denoising chain (BPTT), using different formulations with distinct limitations. Motivated by KL-regularized RL, we unify these seemingly different methods under Generalized Optimal Policy Matching (GOPM), a policy-agnostic framework that casts soft policy improvement as matching policy log-ratios to soft advantages under a general divergence. For diffusion policies, we instantiate GOPM as Diffusion Advantage Matching (DAM), a theoretically grounded method with a detached estimator that avoids BPTT while preserving the soft-optimal distribution. Under the Log-Variance divergence, DAM remains valid for samplers covering both policy and reference, enabling DAM to support online, offline, and hybrid learning under a single objective; to the best of our knowledge, it is the first diffusion-RL method to do so. Across MuJoCo and DeepMind Control, DAM achieves strong online and competitive offline performance, better preserves multimodal distributions, and delivers strong results in both single-stage and two-stage offline-to-online learning on D4RL.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.