acceptodds
Under review as a conference paper at ICLR 2027

Displacement Matching: One Objective for Reward Learning and Distillation in Few-Step Generators

Abstract

Few-step generators synthesize images with a handful of deterministic updates, but post-training them faces two difficulties: the deployed sampler lacks a stochastic transition likelihood for standard policy gradients, and reward alignment and distillation lack a unified formulation. We take a distributional view, expressing both forms of supervision as multiplicative tilts of the generator's terminal distribution. Forward noising associates clean outputs with noisy observations, yielding a clean-sample posterior at each noise level. The terminal tilt shifts these posterior means, defining a displacement in the student's clean-prediction space. We introduce *Displacement Matching* (DM), which learns this posterior displacement as an additive prediction update, unifying pure distillation, sequential reward alignment, and joint training in one regression. DM estimates the generator's posterior mean separately from re-noised terminal samples, since a few-step head represents a finite-interval jump and need not equal that mean. The resulting regression accommodates black-box rewards without requiring trajectory likelihoods or differentiation through the rollout. Our population analysis connects these updates to deterministic deployment, establishing first-order reward improvement up to an explicit finite-step residual and a unique teacher fixed point for raw distillation. At matched reward-query budgets, four-step DM improves alignment and distillation quality while preserving fidelity, outperforming SOTA methods with 0.362 vs. 0.345 HPS on SD3.5-M and 0.93 vs. 0.81 GenEval on Z-Image.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.