MetaDMD: Stable and Scalable Training of Posterior Samplers for Test-Time Steering
Abstract
Test-time steering aligns pretrained diffusion and flow models with arbitrary rewards without fine-tuning, but it requires sampling clean data from the posterior given each intermediate noisy state. Exact posterior sampling via SDE or ODE simulation is computationally expensive, and repeating it at every step makes the cost prohibitive. Stochastic flow maps reduce this cost by learning a flow map that transports Gaussian noise to the posterior in a few steps. However, they rely on consistency objectives that distill the exact integration of this trajectory, which makes training unstable and slow to converge. We observe that steering requires only samples that follow the posterior, not the trajectory that reaches them. Based on this, we propose MetaDMD, which learns a few-step posterior sampler by matching posterior distributions instead of individual trajectories. MetaDMD requires no consistency loss, or pretrained flow map, thus trains stably even on large text-to-image (T2I) models. On class-conditional ImageNet with SiT-XL/2, MetaDMD outperforms stochastic flow maps in test-time steering with reward gradients and performs comparably without them. On SD3.5-M, it also outperforms them across various rewards on DrawBench and on GenEval, while converging approximately faster.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.