acceptodds
Under review as a conference paper at ICLR 2027

From Numerical Rules to Answer Tokens

Abstract

The effect of an intervention on a language model's hidden state can depend on the token used to report an answer. We test this dependence on parity and divisibility by three, exchanging the answer labels while keeping the number and rule fixed. Class directions fitted under the two mappings yield a symmetric component that keeps its sign under the exchange and an antisymmetric component that reverses it. In Qwen2.5-7B, natural projection removal of the symmetric component at layer 22 lowers the correct-answer logit margin by about three, while removing the antisymmetric component changes it by less than 0.1. At the final layer, their order of influence reverses. This change holds under norm-matched edits, across three calibration splits, and for divisibility under a second answer vocabulary. An exact analysis of the final readout explains the final-layer effect. The intermediate symmetric direction also supports classification across answer vocabularies: a classifier fixed on one vocabulary recovers all 100 held-out parity labels under another mapping where the model's own candidate answers are only 50% accurate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.