acceptodds
Under review as a conference paper at ICLR 2027

The Representation-Continuation Gap: Low-Rank States vs. Low-Rank Dynamics in Vision Transformers

Abstract

Low-rank activations make a Vision Transformer look compressible, but compression has to survive the remaining residual computation: after a state has been represented in a small subspace, the later blocks must still compute there. We show that these two requirements can diverge. A subspace can reconstruct the current state, or entry, nearly exactly while failing to support the suffix computation, a failure we call the representation-continuation gap. In eight frozen ViT settings spanning supervised, self-supervised, and CLIP pretraining, confining only the last two blocks to an almost lossless entry space reduces accuracy by 1 to 67 points (median 12). The gap persists under distribution shift and at mid depth, and it appears even though late-layer subspaces are image-specific and stable across views. The obstruction is directional rather than merely low-rank: at a 64-direction budget, replacing eight entry directions with directions read from the input's natural future raises CLIP-FT ViT-B from 3.9% to 82.5%, while random directions reach 4.9%. We explain this with two facts: a fixed continuation space must contain directions created after the entry, and whether an omitted direction matters depends on how strongly the output responds to it. We then introduce functional subspace completion (FSC), a small self-distillation of subspace directions. FSC keeps the dominant entry directions, learns a shared source of complementary directions through the frozen suffix by matching the original model's predictions, and uses no future states at test time. With only 6,144 learned parameters per ViT-B classifier (0.007% of the frozen model), FSC raises accuracy on 5,000 held-out images from 50.6% to 70.5% on AugReg ViT-B and from 4.0% to 77.6% on CLIP-FT ViT-B relative to the entry basis, against 83.0% and 84.5% for unconstrained continuation. FSC improves on the entry basis in all eight settings and transfers to shifted test sets without retraining. Low-rank spaces should therefore be evaluated not only by the states they reconstruct, but by the computation they allow to continue. Code is available in the supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.