acceptodds
Under review as a conference paper at ICLR 2027

Relational Linearity: Probing Language Models Through Their Own Predictions

Abstract

We study relational linearity in language models: whether a single linear map on the context representations alone (“Yesterday, we watched a movie”) reproduces the model own next-token distribution for a fixed query (“What is the tense of theprevious sentence?”) over candidate answers (“past”, “present”, “future”). We introduce LiRe, a probe fitted by matching the logits of the prompted model, thus requiring no ground-truth annotation, and apply it at every layer of four language models on six queries of increasing abstraction. We find that relational linearity is more pronounced in larger models and tracks the abstraction a query demands: surface-level linguistic relations (language, verb tense) are already recoverable in early layers, whereas more abstract ones (subjectivity, faithfulness) peak in middle layers. It is also tied to how the query is formulated: e.g., translating it, affects smaller models markedly and larger ones only mildly. Finally, relational linearity can be high even when the model most likely candidate answer disagrees with the ground-truth label, highlighting a structure that would be missed if probed only with ground-truth annotations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.