Predicting Linear Relational Decodability Without the Relation
Abstract
Linear relational embeddings (LREs) approximate a language model's subject-to-object mapping with an affine function. Their faithfulness varies enormously across relations, from near zero for `person_father` to the ceiling for `country_capital_city`. We investigate the factors that contribute to this variation and demonstrate that it can be significantly predicted from a readout that is not aware of the underlying relation. To do so, we propose a novel form of the Jacobian lens, the *entity lens*. Whereas the Jacobian lens is estimated across all tokens and concepts, the entity lens is trained to predict what the model will output for an entity of type under the neutral prompt ` is known for`; e.g., we train ELs for countries and persons. After estimating an entity lens, we measure how highly it ranks the model's answer for a relation without using any relation-specific information in estimation and prediction. Across 28 relations of the LRE dataset on OLMo-3-7B this relation-agnostic score correlates with LRE faithfulness at (). This suggests that linear decodability is largely a property of what a subject's representation already carries, rather than of the relation being decoded. Moving the unit of analysis from the relation to the individual fact, we conclude that new relation datasets and analysis methods are needed to address several problems in existing datasets, most prominently semantically weak first tokens (e.g., most predicted universities' first token is “University”) and reference competition (e.g., the LLM confuses obscure actor names with competing non-actor entities that have the same last name). Our full pipeline is available [here](https://anonymous.4open.science/r/entity-lens-8CFD/).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.