DMA: Pixel-space Distribution Matching with Adversarial and Anchor Losses
Abstract
Few-step distribution matching distillation (DMD) accelerates latent diffusion, but its latent-tuned recipe transfers imperfectly to pixel teachers. We analyze both sides of DMD. On the teacher-matching side, our diagnostics provide evidence that low-noise RGB matching is dominated by a local-texture cue; this motivates a fixed high-noise matching band. On the real-data side, native clean-RGB outputs allow guidance from an external visual representation without traversing a decoder or sharing the heavy fake-score critic. DINO-Adv removes this critic from the adversarial gradient path and supplies local parametric patch guidance. For distribution-level guidance, we introduce AF-Loss, a parameter-free semantic distribution-field objective for DMD. It turns detached empirical real and generated feature supports into a multi-scale field and injects its stopped displacement as an auxiliary generator target. AF-Loss requires no learned field estimator, adds no learnable parameters, and introduces no inference-time computation. Together these designs form DMA. Across DPG-Bench, GenEval, VQAScore, and COCO30K, the four-step DMA student performs better than the 25-step teacher on all seven reported metrics and ranks first among the evaluated few-step distillers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.