acceptodds
Under review as a conference paper at ICLR 2027

Geometric precursors of learning transitions

Abstract

A model can begin reorganizing internally before a new capability becomes meaningfully detectable in its output behavior. After that point, intervention may require expensive rollbacks during training, so efficient steering requires cheap early detection methods. To this end, we develop a geometric framework for understanding and detecting such changes by tracking latent representations on relevant inputs. By constructing Gram matrices out of latent representations, we are able to give an invariant decomposition of motion during learning into growth, spectral allocation, internal rotation, and support transport. Our central theoretical result relates the amplitude of an emerging response to the rate of its geometric motion, explaining when normalized geometry can change before amplitude-dependent measurements reveal it. Using this description, we derive scalar early warning signals of learning. We then test this method in three settings (deep linear networks, induction heads, and the Pythia suite), each with decreasing prior knowledge about future learning events. In deep linear networks, support transport predicts the next acquired mode. For induction, alarms based on representations alone precede behavioral expression and remain silent on controls. Across Pythia scales, we identify recurring internal reorganizations and a shift from early growth to later internal rotation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.