acceptodds
Under review as a conference paper at ICLR 2027

Reconstruction Does Not Determine Generative Learnability: A Theory of Resource-Matched Visual Tokenization

Abstract

A visual tokenizer is typically optimized for reconstruction, while the downstream generator must learn the latent distribution it induces. We show that better reconstruction need not imply easier generation because reconstruction does not identify the generative geometry of the latent space: reconstruction-equivalent tokenizers can induce sharply different finite-resource learning problems. To quantify this gap, we develop Rate-Distortion-Generative Learnability (RDGL), supported by diffusion-scale representation complexity, latent-to-image transport bounds, finite-sample minimax limits, a full-spectrum deterministic equivalent from random matrix theory, and separate capacity and optimization laws. The resulting theory shows that the preferred representation is resource dependent and motivates Capacity-Matched Spectral Tokenization (CMST), which shapes latent geometry to the downstream data, model-capacity, and optimization budgets. Experiments support both the mechanism and the resulting design principle: reconstruction-preserving latent gauges raise fixed-budget gFID from 4.18 to 15.61, while CMST reduces gFID from 4.18 for the matched VA-VAE control to 2.96 under the same downstream protocol. The resulting principle is that a representation's complexity and geometry should be matched to the learner that must model it, rather than optimized for reconstruction alone. The anonymized code is available at https://anonymous.4open.science/r/ICLR2027-CMST/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.