CINDI: Coarse-to-Fine Implicit Neural Diffusion for Efficient Pixel-Space Generation
Abstract
Pixel-space diffusion and flow-based models typically perform every denoising step at the full image resolution. Although early steps primarily establish coarse structure, sampling them still incurs the full cost of fine-resolution processing. We ask whether spatial resolution can instead be scheduled along the sampling trajectory, allowing early function evaluations to operate on coarser grids while preserving final generation quality. To answer this question, we present CINDI (Coarse-to-fine Implicit Neural DIffusion), a pixel-space generative model that progressively increases spatial resolution along the sampling trajectory. CINDI features an end-to-end implicit transformer denoiser trained with the next-scale bridge loss to enable both in-scale and cross-scale prediction, and a companion coarse-to-fine sampler that operates without the need for multi-model cascading or latent-space compression. We compare CINDI against a single-scale JiT baseline that achieves state-of-the-art-level performance on ImageNet. With matched model size, training epochs, and sampling steps, CINDI achieves competitive FID scores while delivering up to a 3× speedup in inference throughput. Additionally, benefiting from its multiscale prediction capability, CINDI also natively supports generation beyond its training resolution without extra fine-tuning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.