Transition Affinity: A Unified Notion of Geometry for LLMs and Text Corpora
Abstract
Transformer next-token predictions induce an argmax partition of the embedding space, where each cell consists of representations for which a given vocabulary token has maximal logit. We first argue that this representation is a useful way to interpret a transformer. We then show that the geometry of this partition is governed by a phenomenon we call transition affinity: the size of the shared region between two cells is highly correlated to the probability of the two corresponding tokens following one another. This yields a common geometric representation, allowing us to compute correlations between two LLMs, two corpora, or an LLM and a corpus. We conclude by showing how transition affinity can be used to detect that a model has been distilled.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.