A Transport Law for Linear Directions Between Language Models
Abstract
The Linear Representation Hypothesis argues that a concept corresponds to a direction in a model's representation space, while the Platonic Representation Hypothesis posits that independently trained models converge to a shared structure of representations. A central unresolved question is whether these two ideas imply that a behaviorally meaningful direction learned in one model should be recoverable and transferable in another. We address this question in a deliberately narrow setting: the **direction of a successful attack** (prompts for which a model produces a response judged unsafe by a guard) within a single modality, text decoders, and the linear case, where both read-out and transfer are linear. Under class imbalance and guard-dependent labels, we first derive an **admissibility criterion** for statistics used to measure transfer of the attack space. From these admissible statistics, we derive a **transport law**: the correlation between a probe's score in its source model and its carried score in the target model equals the root-mean-square canonical correlation of the model pair, weighted across directions by the probe's profile. We further prove an exact decomposition of transport into within-class transport and label transport. The decomposition reveals a striking asymmetry: labels transport without loss, and on some pairs with a gain (–), whereas the underlying direction is only partially recovered (–). We validate the transport law across 11 dense transformers ranging from B to B parameters, yielding 110 ordered model pairs. Their responses are labelled by ten diverse guards using a corpus of prompts, with representations compared at fourteen relative depths. We further show that the law holds beyond text decoders, on 8 alternative architectures ranging from B to B parameters. In classical decoder transformers we find that transfer is asymmetric in model size, where a probe transferred from a smaller model to a larger model loses markedly less AUC than transfer in the reverse direction, and the difference increases with the size gap. Together, these results provide a principled framework for testing whether a **shared linear direction of successful-attack behavior exists across independently trained models and transfers between them**. More broadly, our work turns two influential representation hypotheses from qualitative claims into testable, quantitative questions about cross-model representation transport, with direct implications for understanding and controlling model behavior in AI safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.