From Representation Correspondence To Computational Trajectories In Language Models
Abstract
Geometric analyses of neural networks characterize correspondences between hidden states, but it is unclear when a correspondence between representations extends to a correspondence between the computations that transform them. We study this question by distinguishing a state correspondence at one depth from the transport of a trajectory by a depth-indexed family of correspondences. Across equivalent number representations, affine low-rank correspondences accurately transport hidden states but fail when applied statically to local residual-stream updates. Transporting the two endpoints of an update with the correspondence fitted at each endpoint instead raises update alignment from to , approaching an update-specific correspondence at . The same matched-sample pattern replicates across musical-pitch notation systems (, versus for an update-specific map). In numbers, decomposing the dynamic predictor identifies evolution of the state correspondence across adjacent layers as the dominant additional component. Independent adjacent-map estimation reduces the effect but retains of the static-to-update-specific gap, while shuffled correspondences collapse to near zero. We then ask whether depth-evolving state correspondences predict computation outside the examples used to fit them. In controlled addition, correspondences learned only from baseline states of training operands predict arithmetic-specific trajectories for held-out operands. A strong late-layer correspondence emerges between Question and Compute formulations, and transported Question–Compute states induce aligned arithmetic-specific updates when executed by the actual destination transformer block. Finally, a semantic-shuffle control separates native computational correspondence from local dynamical compatibility: shuffling identity pairings destroys state alignment, native trajectory prediction, and interventional update recovery, while substantial layerwise square closure remains. Thus a representational snapshot does not generally specify a computational trajectory, but the evolution of representation correspondences across depth can contain predictive information about that trajectory. Computational correspondence should therefore be validated against native computation rather than inferred from state alignment or local dynamical compatibility alone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.