From Numerical Rules to Answer Tokens
Abstract
The effect of an intervention on a language model's hidden state can depend on the token used to report an answer. We test this dependence on parity and divisibility by three, exchanging the answer labels while keeping the number and rule fixed. Class directions fitted under the two mappings yield a symmetric component that keeps its sign under the exchange and an antisymmetric component that reverses it. In Qwen2.5-7B, natural projection removal of the symmetric component at layer 22 lowers the correct-answer logit margin by about three, while removing the antisymmetric component changes it by less than 0.1. At the final layer, their order of influence reverses. This change holds under norm-matched edits, across three calibration splits, and for divisibility under a second answer vocabulary. An exact analysis of the final readout explains the final-layer effect. The intermediate symmetric direction also supports classification across answer vocabularies: a classifier fixed on one vocabulary recovers all 100 held-out parity labels under another mapping where the model's own candidate answers are only 50% accurate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.