The Deviation Carries the Code: Causal Surgery on Geometric Representations in Language Models
Abstract
Causal analyses locate language-model arithmetic in geometric structures like the month circle and the number helix. Each characterizes the representation through that structure. We call the remainder orthogonal to that structure the deviation. Recent theoretical work predicts such components exist, but no work has manipulated their content. We test whether models rely on them by editing token embeddings, rewriting only tested components. A matched-norm control substitutes random content, isolating content effects from magnitude effects. Across four open-weight models and a grokked transformer, on month, weekday, hour, letter, and two-digit addition arithmetic, the structure alone never restores the unedited model's accuracy. In month, hour, and letter arithmetic, the deviation alone restores 70 to 102% of that accuracy. Weekday arithmetic varies by model, and addition requires both components. Rotating the circle one step opposes the deviation, which we shrink to half its norm. In Llama-3.1-8B, the answer follows the deviation on 95% of prompts and the circle on 1%. Patched at full norm into three models' residual streams, the deviation wins month and weekday conflicts at every tested layer where the operand position still controls the answer. For detecting such structure, strength-based readouts are uncalibrated, reporting structure in 93.5% of structureless inputs matched to the GPT-2 embedding spectrum, whereas our significance test does so for only 4.0% of those inputs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.