acceptodds
Under review as a conference paper at ICLR 2027

PatchBridge: Decoupling Patch Roles for Pixel-Space Diffusion Transformers

Abstract

Pixel-space diffusion Transformers denoise RGB images directly, avoiding the fixed autoencoder compression interface common in latent diffusion. Yet many process one patch grid throughout the backbone, asking it to support global context modeling, full-depth denoising, and late local refinement. These roles differ in their spatial and computational requirements; we call their forced sharing of one grid patch-role coupling. We introduce PatchBridge, which assigns role-specific patch sizes and lifetimes to three coordinated streams. A compact semantic stream maintains global context through sparse updates, a structural stream carries the principal denoising state through the full backbone, and a finer detail stream activates in later layers for local refinement. This allocation preserves full-depth denoising without processing the finest grid at every layer. Parent-to-child spatial modulation supplies aligned context to finer streams, while child-to-parent pooled feedback returns updated evidence to coarser streams, keeping their representations coordinated as views of one evolving denoising state. At the output, an overlapping clean-image readout decouples detail-token stride from prediction support, allowing neighboring tokens to jointly estimate shared pixels. PatchBridge achieves FIDs of 1.44 and 1.54 on class-conditional ImageNet at 256x256 and 512x512, respectively. In text-to-image generation, it scores 0.87 on GenEval and 84.1 on DPG-Bench. These results support role-specific patch allocation for direct pixel-space generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.