acceptodds
Under review as a conference paper at ICLR 2027

DIFFUSIONRVGE: REINFORCEMENT LEARNING VIA REGRESSION-BASED VALUE GRADIENT ESTIMATION

Abstract

Reinforcement learning improves diffusion models using rewards on generated images, but repeated denoising and gradient estimation make training expensive. We introduce DiffusionRVGE, an efficient method that estimates useful update directions by regressing centered rewards on centered final denoised images. It requires neither reward derivatives nor adjoint backpropagation through the sampling trajectory. A unified first-variation formulation gives a gradient-matching update with or without a trajectory-KL penalty, and clean-image regression supplies a preconditioned approximation to its value-gradient target. A supporting analysis explains the directional advantage of clean-image over noisy-state regression. The implementation adapts rollout bandwidth and query weights to reward continuity: local rollouts suit smooth rewards, whereas broader groups provide informative contrasts for discrete rewards. Trained jointly on five rewards, one SD3.5-Medium model matches or exceeds DiffusionNFT's joint model on seven of eight metrics, reaching GenEval 0.94, OCR 0.92, and PickScore 23.82 at about one fifth of its estimated training compute. In single-reward training, Diffusion-RVGE reaches matched reward thresholds 2.1-6.7x faster than reproduced DiffusionNFT on identical hardware. Ablations support the clean-image estimator, direct direction transfer, and reward-dependent rollout choices. Our code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.