AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer
Abstract
Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoising remains underexplored. In this paper, we define the semantic recovery objective: the denoising process should recover the semantic content of the ground-truth image from noisy latent, and a good tokenizer should make it easier. Existing approaches train a projector to predict the semantics directly from the noisy latent. We argue that this predicts the average of clean-image semantics, whereas what really needs to be aligned is the semantic of averaged clean latents. More importantly, we demonstrate that the semantic recovery error orthogonally decomposes into the error of the optimal semantic prediction directly from the noisy latent and the error between these two semantic predictions. We therefore identify their consistency as the missing requirement and call it Semantic Affine Consistency (SAC). To examine whether this overlooked requirement is closely related to downstream generation, we introduce Msac, a tokenizer-side proxy for SAC. Across the evaluated tokenizers and diffusion model scales, Msac closely tracks generation quality, reaching a Pearson correlation of with SiT-XL gFID, thereby establishing SAC as a practical diagnostic and motivating SAC-guided tokenizer training. We then introduce AffineTok, which promotes SAC through two complementary, training-only components. Global Semantic Coordination Token(GSCT) coordinates the semantic organization of clean latents, keeping semantic averaging meaningful, while Posterior-Mean Semantic Alignmen (PMSA) predicts posterior-mean latents from noisy inputs and supervises their semantics. On ImageNet \(256\), compared with the baseline, AffineTok reduces gFID by \(26%\) at 20 epochs and, with continued training, achieves a new state-of-the-art gFID of \(1.21\) without classifier-free guidance and 1.10 with guidance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.