Representation Without Capability, Then Capability Without Abstraction: Parity in GPT-2
Abstract
Probing has become a standard instrument for characterizing what a language model represents, and it underlies a family of intervention methods, among them directional ablation and activation steering, that manipulate behavior aling probe-derived directions. These methods operationalize a rarely tested assumption: that the direction along which a concept is linearly decodable (representation) is the direction the model reads and uses (capability). Evaluating this presupposition is non-trivial for semantically rich concepts because a robust test requires an independent ground truth about what the model actually computes. We therefore construct a setting in which that ground truth is available and the rule the model implements can be identified from behaviour. We use the pretrained GPT-2 small model, whose representations were formed by ordinary language modeling, a natural setting. We observed that pretrained GPT-2 cannot perform parity detection above chance in any prompt format we administered, while a hyperplane in its embedding space separates numerals by parity at 0.994 on held-out tokens: Representation Without Capability. This representation the model cannot act upon is selective, as divisibility by three performs at 0.687. We then fine-tune it on a three-operand arithmetic verification task adapted from an intracranial-EEG protocol for human subjects kalinova2026temporal: with false candidates at . We fine-tune on both the easy and difficult variants, which differ in operand and result digit-width: single-digit operands with results in , versus a two-digit with results in . Arithmetic is not an intrinsic capability at this model scale and remains heuristic at far larger ones. We observed GPT-2 constructing a chained-parity solution: the fine-tuned model attains 97.9–100% accuracy by comparing the candidate result's parity against . Neither constituent of this solution originates in fine-tuning: parity is already separable in the embeddings, and the pretrained model already computes XOR at 0.91–1.00 over other feature-disjoint pairs. Fine-tuning enables the connection between parity identity and XOR, turning on the chained-parity detection capacity that solves the task. But the Capability it creates is not Abstract: behavioral tests suggest that the model stores a parity bit per token-position pair but does not compute parity from the numeral. A token never seen in contributes always even, and an odd three-digit and the word “cat” become indistinguishable, both at 0.985 and mean 8.25. XOR composition was inherited from pretraining. Causal ablation confirms parity is used upon fine-tuning in the arabic easy task, as removing it costs 6.5 accuracy points at rank 40 against 0.2 for an equally decodable property of the same token and 0.0 for random directions. In the difficult condition, however, parity and the matched control damage the model equally at rank 80 (0.9554 vs 0.9582, against a 0.9826 baseline) while random directions do not (0.9813), suggesting that parity over a ninety-token operand occupies far higher rank than over a ten-token one. Interestingly, learnings from the easy condition show that removing parity leaves the model confident and wrong while random ablation leaves it hesitant and right. The same signature appears when we break a token slot's surface format, an intervention applied at the input that touches no activations. Making parity usable also dispersed it: an identical 40-dimensional cut reduces its decodability to 0.4916 in the pretrained model—chance—but only to 0.8183 after fine-tuning. Our findings have practical implications for questioning how the representation-capability relation is tested. Standard procedures of directional ablation, computing the direction on a forward pass and projecting it out finds nothing: accuracy stays at 1.000, failing because every block writes parity back into the residual stream along fresh directions. Altogether, this paper suggests that representation does not entail capability, and capability does not guarantee abstraction. These observations have consequences for the literatures of probing and benchmarking, and for current views on concept erasure and approaches for unlearning. Beyond, we connect our findings with safety discussions, since a procedure that activates dormant knowledge is powerful when intended and hazardous when it arises as an unanticipated side effect of postraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.