acceptodds
Under review as a conference paper at ICLR 2027

Aligning Latent Geometry for Spherical Flow Matching in Image Generation

Abstract

Latent flow matching typically treats pretrained variational autoencoder (VAE) representations as fixed, although accurate reconstruction does not ensure that their geometry is well suited to downstream generative modeling. We study whether adapting the latent geometry improves image generation while retaining the encoder and diffusion architecture. Across three tokenizer families, component-swap probes show that preserving token direction retains substantially more decoded perceptual and semantic similarity than preserving radius. Based on this observation, we project each latent token onto a fixed-radius sphere, with the encoder frozen, and lightly finetune the pretrained decoder to reconstruct from the projected representation. We then train spherical flow matching on the induced spheres. On class-conditional ImageNet-256 with FLUX.2 SiT-B/2, the complete method reduces FID from to under a matched training recipe and a fixed 50-update sampling protocol, and reaches FID in approximately fewer flow-training steps. Matched decoder controls support the combined latent construction and adaptation. Recipe-transfer comparisons yield lower FID across three tested tokenizer families and SiT-B/XL backbones, and the gain remains under a matched representation-alignment objective. Additional controls characterize reconstruction across the evaluated tokenizers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.