DASH: Steering Attention Heads for Fine-Grained Concept Control
Abstract
Attention heads are widely used as units of analysis and intervention in transformer models, but treating an entire head as concept-specific can make interventions too coarse. We study concept control at the level of head–direction pairs and introduce DASH, Direction-Aware Steering of Heads. Given a concept dictionary and a selected set of attention heads, DASH identifies concept-related directions within each head and selectively modifies their activation components without retraining the model. Experiments on toxicity suppression, image captioning, image classification, and factual question answering show improved control–preservation trade-offs over whole-head rescaling across a broad range of settings. We further demonstrate simultaneous editing of multiple attributes, including when their selected heads overlap, while largely preserving an untargeted attribute. Together, these results demonstrate the value of within-head directional editing for selective and composable concept control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.