acceptodds
Under review as a conference paper at ICLR 2027

AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer

Abstract

Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoising remains underexplored. In this paper, we define the semantic recovery objective: the denoising process should recover the semantic content of the ground-truth image from noisy latent, and a good tokenizer should make it easier. Existing approaches train a projector to predict the semantics directly from the noisy latent. We argue that this predicts the average of clean-image semantics, whereas what really needs to be aligned is the semantic of averaged clean latents. More importantly, we demonstrate that the semantic recovery error orthogonally decomposes into the error of the optimal semantic prediction directly from the noisy latent and the error between these two semantic predictions. We therefore identify their consistency as the missing requirement and call it Semantic Affine Consistency (SAC). To examine whether this overlooked requirement is closely related to downstream generation, we introduce Msac, a tokenizer-side proxy for SAC. Across the evaluated tokenizers and diffusion model scales, Msac closely tracks generation quality, reaching a Pearson correlation of with SiT-XL gFID, thereby establishing SAC as a practical diagnostic and motivating SAC-guided tokenizer training. We then introduce AffineTok, which promotes SAC through two complementary, training-only components. Global Semantic Coordination Token(GSCT) coordinates the semantic organization of clean latents, keeping semantic averaging meaningful, while Posterior-Mean Semantic Alignmen (PMSA) predicts posterior-mean latents from noisy inputs and supervises their semantics. On ImageNet \(256\), compared with the baseline, AffineTok reduces gFID by \(26%\) at 20 epochs and, with continued training, achieves a new state-of-the-art gFID of \(1.21\) without classifier-free guidance and 1.10 with guidance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.