acceptodds
Under review as a conference paper at ICLR 2027

ROAD: Modern Adversarial Training for Simple One-Step Distillation at Scale

Abstract

One-step distillation turns pretrained diffusion models into fast generators, but strong results often require substantial training machinery and compute. Pure adversarial distillation offers a simpler route, yet has struggled at modern text-to-image scale. To address this limitation, we propose ROAD (Representation-grounded One-step Adversarial Distillation), a one-stage framework that distills text-to-image models with one adversarial objective and a lightweight representation-space critic. First, we replace the classical convolutional head with a Lipschitz-controlled Transformer head with token-level text conditioning, which matches the semantic tokens of modern visual encoders. Second, we observe that fixed random feature mixing leads to long-horizon degradation. We attribute this to generator–critic co-adaptation around one realization-specific view and periodically re-draw the mixing to keep the randomization effective. Third, for our dense and vector-valued logits, we introduce Jacobian R1 which penalizes the full input–logit Jacobian, overcoming pooled R1's limitation of neglecting most logit sensitivities. A single adversarial objective then suffices: in one step, ROAD reaches 11.03 FID on SDXL and 13.73 on SD3.5-M on COCO-10K, surpassing the multi-step teachers at a fraction of distribution matching's training compute.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.