Concept Geometry in Language Models: Cones, Edits, and Compositional Paths
Abstract
A language model is commonly thought to represent a concept as a direction in its hidden states, found by averaging how the state changes when one word of a sentence is swapped. The average hides whether the individual changes agree, and it says nothing about several words changing at once. We keep every change. Repeating a counterfactual edit, such as a change of tense or a synonym for an adjective, across contexts and across the words that carry it gives a cloud of vectors shaped like a cone, with an axis and a width; comparing the axes of different words for the same edit with those of random words in the same position tells whether the model treats the edit as one change or as several. Because an edit's axis barely depends on the sentence it acts on, the sentences reachable by several edits form a lattice, and their combined effect divides among the edits by the Shapley value, each edit's contribution averaged over the orders in which the edits can be applied. The lattice grows exponentially with the number of edits, yet a single path that applies the edits one at a time already estimates this division: apart from sampling noise, its error is the interaction between edits, which averages out over a few paths. Across five models from three families, synonyms of an adjective scatter more than random adjectives in all combinations of model and sentence, while verbs of destruction align more than random verbs in . Shares vary little with the order of the edits (standard deviation at most over all orders of six edits). On a lattice of sentences, at most eight paths, visiting of the sentences, recover every share to within , and the shares predict which edit's removal most weakens a combined steering vector (Spearman ).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.