acceptodds
Under review as a conference paper at ICLR 2027

Where Does a Representation's Value Live? The Realizability Decomposition of Dimensionality Reduction and Depth

Abstract

Methods grouped under dimensionality reduction are used to obtain lowerdimensional descriptions of data, but “fewer coordinates” does not specify the geometric operation that produced them. Against the reference of a basis change, which keeps all D coordinates, we separate two constructions that both return k < D numbers: a projection onto a subspace of the original ambient space, and an embedding into a distinct target space. Every rank-k linear map is of the first kind, an orthogonal projection followed by an invertible coordinate change with a closed-form metric that restores the projected geometry, so the distinction has content only for nonlinear and learned maps. To measure it we introduce the realizability decomposition: any finite centered representation Z of a declared source X splits orthogonally as Z = Z∥ + Z⊥, where Z∥ is exactly the part expressible as coordinates of one fixed source-subspace projection and Z⊥ is what no such projection can produce. The residual ρ = ∥Z⊥∥ / ∥Z∥ comes with an exact finite-sample caveat and a random-representation null. The decomposition also makes precise what a linear autoencoder identifies (a subspace and a metric, never a basis) and why a decoder trained with a pointwise loss returns to the original space rather than a symmetric copy of it. Empirically, a fixed projection of pixel space carries 82–95% of the energy of t-SNE, UMAP and PaCMAP layouts of MNIST and Fashion-MNIST, and 91–95% on CIFAR-10, yet that part has only PCA-level neighborhood fidelity; the embeddings’ advantage lives in the small remainder Z⊥. For autoencoders, end-to-end training produces codes with Z⊥ ̸= 0 whenever it beats PCA, but a nonlinear decoder trained on PCA scores (Z⊥ = 0) beats PCA as well, so the remainder describes the code training chose, not a requirement of reconstruction. The decomposition is a diagnostic, not a theory of representation: it says how much of a representation is a projection of a declared source, and what follows from that.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.