Graph Representations in J-Space
Abstract
Language models appear to have a global workspace: a small part of their internal state that holds active concepts and supports multi-step reasoning. The Jacobian lens reads this workspace as individual words, but reasoning also requires knowing how things are related. Answering “Who is the teacher of the banker of John?” requires following a graph whose nodes are people and whose edges are relations. We investigate whether the workspace holds these edges or only the nodes. We compare questions about the same story that differ in one relation word, and edit the workspace to test which changes affect the answer. Across multiple models, the workspace holds both nodes and edges, but treats them differently. Editing a person affects the answer where the name appears and at every model depth; editing a relation matters only at the relation word and in the middle layers. The two are never available at the same place and time, so the workspace never holds the connected graph. By the time the answer forms, the relation that determined it has left the workspace, although it remains elsewhere in the model. Workspace inspection at the moment of answering can therefore miss the relation responsible for the answer. Our findings motivate interpretability methods that trace relational information across token positions and layers, rather than relying on a final workspace snapshot.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.