SHADOW ANCHORING: GROUNDING HUMAN- INTERPRETABLE CONCEPTS IN TRANSFORMERS VIA SPARSE AUTOENCODERS
Abstract
Interpretability methods for language models are mostly post hoc: probes and sparse autoencoders (SAEs) are fit after training to discover features the model happened to learn. We study a proactive alternative, concept anchoring, in which designated latent coordinates are trained to carry human-specified concepts, and ask whether it can be retrofitted into a pretrained transformer without degrading language modeling (an "interpretability tax"). We use continuous RGB color in CSS code as a machine-checkable testbed, since hex literals provide exact three-dimensional ground truth inside natural, syntactically constrained text. We introduce SAE shadow anchoring: supervision is applied to three latents of a frozen, pretrained SAE at an intermediate layer of Qwen2.5-Coder-0.5B, and gradients flow through the frozen SAE into LoRA adapters. Compared with anchoring the raw residual stream, shadow anchoring reaches higher color recovery with a smaller language-modeling cost relative to unconstrained LoRA fine-tuning. Clamping the anchored latents shifts generated color values in a dose-dependent way, keeps generated CSS syntactically valid, and introduces no spurious colors on general Python and JavaScript prompts. The representation transfers across six syntactic contexts, and Hewitt-Liang control tasks indicate the probes measure geometry rather than lookup shortcuts. Separately, digit-pair tokenization lets a 38M-parameter transformer trained from scratch generalize to unseen hex codes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.