acceptodds
Under review as a conference paper at ICLR 2027

Rethinking the Design Space of Vanilla Pixel Diffusion Transformers

Abstract

Pixel-space diffusion models have recently re-emerged as a compelling alternative to latent diffusion, avoiding the lossy latent bottleneck and separately trained autoencoder. JiT shows that a vanilla diffusion transformer can generate directly in pixel space by replacing the commonly used velocity prediction () with clean image prediction (). In this work, we revisit this choice and introduce **Mꜰʟᴏᴡ**, an adaptive parameterization that learns to interpolate between and prediction across principal directions of the data, which can be learned end-to-end with diffusion model training. We find that the preferred parameterization varies systematically with the data representation and model capacity. Notably, increasing model capacity shifts preference toward prediction from $x_0$ prediction. We evaluate Mꜰʟᴏᴡ across ImageNet, UCF-101 video, and large-scale text-to-image generation, demonstrating competitive or improved generative performance across diverse settings. In addition, we find that vanilla pixel diffusion transformers, irrespective of parameterization, retain a persistent high-frequency deficit relative to real images. To address this, we introduce **_ConvRefiner_**, a lightweight convolutional module trained post-hoc with adversarial loss while keeping the diffusion backbone frozen. With only a small amount of additional training, _ConvRefiner_ restores the high-frequency statistics of generated images while preserving FID.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.