acceptodds
Under review as a conference paper at ICLR 2027

Diffusion Advantage Matching: Unified Diffusion Policy Optimization

Abstract

Diffusion policies can represent multimodal behaviors, but improving them online is challenging because their action likelihoods are intractable. Existing diffusion-RL methods address this difficulty through action gradients, likelihood approximations, or backpropagation through the denoising chain (BPTT), using different formulations with distinct limitations. Motivated by KL-regularized RL, we unify these seemingly different methods under Generalized Optimal Policy Matching (GOPM), a policy-agnostic framework that casts soft policy improvement as matching policy log-ratios to soft advantages under a general divergence. For diffusion policies, we instantiate GOPM as Diffusion Advantage Matching (DAM), a theoretically grounded method with a detached estimator that avoids BPTT while preserving the soft-optimal distribution. Under the Log-Variance divergence, DAM remains valid for samplers covering both policy and reference, enabling DAM to support online, offline, and hybrid learning under a single objective; to the best of our knowledge, it is the first diffusion-RL method to do so. Across MuJoCo and DeepMind Control, DAM achieves strong online and competitive offline performance, better preserves multimodal distributions, and delivers strong results in both single-stage and two-stage offline-to-online learning on D4RL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.