The Rank Bottleneck in Diffusion Transformers: Why and When Width Must Track Token Dimension
Abstract
Diffusion transformers on high-dimensional latent tokens have been reported to require a width comparable to the token dimension, even when fitting a single training example. We show that this requirement can arise from a rank bottleneck in the final per-token linear layer, rather than from a limitation of the transformer backbone itself. A width-d output layer can emit only a rank-d subspace of an n- dimensional token. When the regression target contains an isotropic component, that component cannot be represented through the layer, giving an error floor of 1 − d/n; more generally, the floor is determined by the target covariance beyond rank d. With nothing fitted, this floor predicts the reported failures with a mean error of 5%. The same bottleneck persists during sampling: a rank-d velocity field cannot alter the component of a sample outside its output subspace, leaving the corresponding part of the initial noise unchanged. This also explains why EDM preconditioning does not remove the floor. In contrast, predicting the clean token and reconstructing the isotropic component from the network input routes that component around the rank-limited output layer. For the clean target, the remain- ing floor can be bounded directly from the target spectrum, giving a criterion for when the ambient token dimension is or is not the relevant width scale. Across latent and pixel representations, our experiments support this view: the observed width requirement follows the effective rank of the emitted target rather than its ambient dimension. These results identify the output parameterization, rather than transformer capacity alone, as the source of the reported width constraint.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.