Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
Abstract
Pre-trained Vision Transformer (ViT) representations are rich and broadly useful, but their high dimensionality makes them expensive to store, transfer, and reuse. We study how to compress these representations while preserving downstream utility, and identify a fundamental geometric obstacle: layer-normalized ViT features concentrate on hyperspherical manifolds, whereas standard Gaussian variational autoencoders model latents in Euclidean space. This mismatch wastes capacity on a radial dimension that carries no semantic information and induces posterior collapse under strong compression. We propose SVAE, a geometry-aware variational autoencoder that replaces the Gaussian bottleneck with a product of Power Spherical distributions, aligning the latent topology with the data manifold while remaining numerically stable in high dimensions. Across five transformer backbones spanning distinct training objectives and modalities, including DINOv2, CLIP, DUSt3R, VGGT, and PAGE-4D, S²VAE consistently outperforms Gaussian baselines under aggressive compression (up to 512×), with the gap widening as the bottleneck tightens. Beyond reconstruction fidelity, we show that the learned latent space is directly usable for downstream tasks. Lightweight decoders with fewer than 0.5M parameters can recover task-relevant signals from compressed latents, and generative models trained in the latent space produce outputs that remain structurally consistent across modalities. These results indicate that geometry-aligned latent design enables compact representations without sacrificing downstream utility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.