acceptodds
Under review as a conference paper at ICLR 2027

Can Capability Enhancement Be Localized?

Abstract

Transformer-based language models have shown strong capabilities across many tasks and domains. However, prior work shows that, for a given task, not all attention heads matter equally. Specifically, they suggest that many redundant heads can be safely removed with little-to-no effect, while a small set of important heads, when ablated, severely damages the model's performance on the given task. In this work, we investigate a complementary aspect, called Capability Enhancement Localization (CEL): given a model and a task, does there exist a small set of heads whose removal improves the task performance?} We first investigate this question empirically across models from different families. Using a simple two-stage routine that screens candidate heads and evaluates their joint ablations, we identify sparse head combinations whose removal improves task-specific performance without further fine-tuning. Motivated by these observations, we ask theoretically whether such a phenomenon can naturally arise from training dynamics, and under what conditions. To this end, we construct a simplified multi-head softmax attention model and two tasks with conflicting features, and show that gradient descent from small random initialization provably concentrates the conflicting features into particular heads. We further show that ablating such heads improves performance on one task while harming the other, even though training converges to a global optimum of the joint objective. This establishes a possible training-induced theoretical explanation for CEL.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.