acceptodds
Under review as a conference paper at ICLR 2027

PixelDDT: Depth-Routed Decoding for Pixel Diffusion

Abstract

Pixel-space diffusion generates images directly in pixel space without a learned image autoencoder. Recent decoupled architectures pair a global transformer with a lightweight pixel decoder, aiming to separate global semantic modeling from local detail rendering. However, pixel-decoder architectures remain insufficiently compared under a shared implementation. We first compare representative decoder families in a shared codebase and identify a compact multiscale U-Net as a competitive design. This architectural split, however, does not determine how the two modules share information. The pixel-decoder receives backbone information only through the final transformer state at its bottleneck. Probes show that this state retains coarse and fine image content, while class separation declines across later backbone blocks. Earlier states, on the other hand, already expose visual content with stronger class separation, motivating direct access to these features at different decoder scales. We therefore introduce PixelDDT, which routes distinct backbone depths to successive U-Net stages while retaining the final-state bottleneck. These connections allow local multiscale reconstruction to draw on earlier features without passing through the terminal backbone state. With matched backbone and decoder widths, depth routing improves ImageNet-256 FID from 2.74 to 2.31 at 400K updates. Gated window attention further exchanges features across neighboring patches, bringing FID to 2.27 at the same training budget. With longer training, PixelDDT achieves FID 1.63 after approximately 380 epochs with classifier-free guidance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.