SemPix: Learning Semantic Flows for Pixel Generation
Abstract
Pixel diffusion generates images directly in pixel space without a separately trained autoencoder. Yet a single pixel trajectory must model object structure and spatial layout together with the diverse textures, colors, and illumination that realize them. Pretrained visual representations reduce sensitivity to such appearance variation and organize images by content and spatial structure, providing a more direct target for the visual structure required by pixel generation. We introduce SemPix, a coupled semantic–pixel flow framework that models pretrained visual representations as an explicit generative state for native pixel synthesis, rather than as auxiliary supervision alone. A semantic flow learns the distribution of dense pretrained features, while a pixel flow generates the complete image from independent noise under spatially aligned modulation by the generated state. This coupling makes semantic structure part of the generative process while preserving stochastic generation in native pixel space. The two flows are trained end-to-end through predicted semantic states, combining representation supervision with pixel-generation gradients that adapt the semantic state for image synthesis. On class-conditional ImageNet 256×256, SemPix converges substantially faster than native pixel diffusion baselines without guidance and achieves competitive generation quality against state-of-the-art pixel generators after full training. These results show that modeling semantic structure as an explicit generative state can reduce the optimization burden of native pixel diffusion.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.