The Orthogonality Gap: Predicting, Before Training, When a Representation (or Foundation Model) Will Help
Abstract
Teams routinely fine-tune a foundation model onto a task only to find it beats no linear baseline—after paying the compute. We show that outcome is decidable in advance, from geometry, for free. Given inputs X already in hand and a candidate representation f, only the part of f orthogonal to span(X) can improve an out-of-sample prediction; a re-encoding of X—which a foundation model computed from X often is—has no such part. We formalize this as the Orthogonality Gap and prove an exact decomposition of the attainable OOS gain (Frisch–Waugh–Lovell composed with an attenuation bound), ΔR²_OOS = g⋆·r²·ρ² ≤ g⋆·r², separating a task-intrinsic ceiling—orthogonal signal g⋆ times target reliability r²—from a representation's alignment ρ² with that signal, with equality characterized. The bound is computed from features alone, with no downstream training, and predicts whether a representation will help and by how much before a single fine-tuning run; empirically it is never violated by a linear downstream and holds in 97% of MLP cases. On four real open-weight foundation models across four domains the forecast is calibrated in both directions: a sentence encoder's positive gain to within 0.01; a frozen single-cell transformer's near-zero gain over linear PCs of its own input; and a protein language model's positive gain on designed-protein stability and enzyme engineering (direction a priori, magnitude conservatively under-predicted)—where acting on the forecast selects designs +0.200 more stable (p=0.002). Applied inside a transformer, the Gap selects the best-transferring layer training-free across four models at 0.004 mean accuracy regret versus probe-every-layer. Against LogME, the strongest training-free transferability score, it is statistically tied on real LLM embeddings while adding a proven ceiling, a reliability term, an incremental form, and a faster variant. We report every negative—a refuted phase transition, a retracted ranking claim.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.