Norm Share Predicts Whether a Subspace Ablation Removes the Feature or the Model
Abstract
Feature ablation is used to isolate what an internal representation contributes to a language model. We show that two operations commonly called ablation produce widely different effects on the same selected features, ranging from language-specific control to broad model failure. For sparse-autoencoder features, latent zeroing subtracts the selected latents' contributions and preserves reconstruction error while directional projection removes every residual-stream component along their decoder directions at every token. Because a language-specific feature's decoder span can carry shared computation, projection can cause broad model failure. We predict this failure using norm share, ρ, the average fraction of a token's squared residual-stream norm in the ablated subspace. This requires one forward pass without labels, generation or intervention. All 52 measured subspaces with ρ above one half raise the other languages' cross-entropy by at least 6 nats. In language-identity tests across eleven languages of parallel text, Gemma 3's mid-network subspaces carry most of a typical token's norm. At layer 17, projection raises target cross-entropy by 7–31 nats (up to several times baseline), with similarly large increases on the other ten languages. Zeroing the same latents concentrates changes on the target and switches output language, with no switches under matched random-feature controls. The same subspace construction at the Llama 3.1 site has low ρ and selective projection, consistent with the decline in projection selectivity as ρ rises across measured subspaces in both families. Manipulating the dominant shared directions changes this outcome: removing them from Gemma's basis makes projection selective, and adding them to Llama's basis lowers selectivity. Adding a random direction of the same rank leaves projection selectivity unchanged. The diagnostic also separates destructive from harmless refusal interventions. We then use selective language removal to test whether language-specific representation interferes with multilingual reasoning. Across three model families and four removal constructions, no setting significantly raises pooled MGSM accuracy. The gains reported by Zhao et al. (2025) are per-language best points of a grid searched on the evaluated problems. Applying this rule to our reproduction, where no setting raises pooled accuracy, gives +2.1 points in-sample and 0.0 when selection and scoring use different halves of the problems. Subspace interventions should therefore report ρ and collateral effects before attributing behavioural changes to removal of the targeted property.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.