acceptodds
Under review as a conference paper at ICLR 2027

Ordinal Axes in the Geometry of Ordered Categorical Knowledge in LLMs

Abstract

Nothing in the word viscount says that it outranks baron, yet language models answer most questions about such ordered categories correctly, even though their order is a convention that cannot be read off the words. Current accounts of concept geometry describe continuous, cyclic and unordered categorical concepts, but not how such orders are represented. We find that ordered categories are organised along an ordinal axis in the residual stream. The axis can be recovered from the activations of the two endpoint categories alone, and it places the held-out intermediate categories in their correct order across several language models and many domains, from ranks and titles to physical scales. We then show that a model uses this axis when it compares two categories. Exchanging their coordinates along this single direction, and nothing else, flips about a third of the decisions it gets right, whereas random directions of the same size flip far fewer. Finally, we show why a trained causal direction outperforms the ordinal axis. A one-dimensional direction trained to flip the comparison combines the ordinal representation with the gradient of the answer logits. Adding this gradient to the ordinal axis recovers much of the trained direction's effect. These results identify a geometric representation of ordered categorical knowledge, read and used by the model, and add ordered categories to the account of concept geometry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.