acceptodds
Under review as a conference paper at ICLR 2027

Semantics as Source: Generative Flows from Semantic Embeddings Instead of Noise

Abstract

Pretrained semantic representations have become increasingly important in modern image generators, serving as conditioning signals, as feature-alignment targets, and even as the latent space itself. However, in every case the generator still begins from a content-free input, typically Gaussian noise. We propose semantics-as-source (SAS), which directly uses semantic embeddings as the source of a generative flow. The flow starts from the image's pooled semantic embedding, perturbed with noise, and the model learns the spatial residual that recovers the image's full latent token grid, which a frozen decoder maps to pixels. At inference, a small prior over embeddings supplies the starting point. Compared with conditioning and alignment on the same backbone and training budget, SAS gives better images and faster training across model scales and conditional and unconditional settings. On ImageNet-1k with SiT-XL/1, our source formulation achieves an unguided FID of 2.71 after only 40 epochs, compared with 2.96 for conditioning and 4.12 for alignment. SAS also roughly halves the expected squared distance the flow must cover, and the source route passes the Gaussian source's final 40-epoch FID within 10 epochs. Swapping only the source into RAE's tuned DiT-DH-XL recipe passes their released 80-epoch model by epoch 50 and beats it at an identical budget; with only 8 sampling steps our model matches the FID RAE reaches with 50 steps. With a decoupled head and autoguidance, SAS reaches FID 1.86 on ImageNet-1k after 160 epochs. These results suggest that semantic information is most effective not merely as guidance for a generative model, but as a structured starting point from which generation proceeds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.