When One Direction Is Not Enough: GRASP for Conditional Activation Steering
Abstract
Activation steering provides a lightweight way to control LLMs without updating model parameters. However, most existing methods represent a target behavior with a single global direction, implicitly assuming that the same intervention is suitable for every input associated with that behavior. Analysis shows that behavior gradients can be low-dimensional yet multi-directional: different inputs may require distinct, sometimes opposing, steering directions. Averaging these directions into a single global vector may therefore miss useful control signals. To preserve this input-dependent structure, we propose Gradient-Routed Activation Subspace Prototypes (GRASP), a conditional activation steering framework that replaces a single global direction with a small set of representative directions. GRASP first projects calibration gradients into a low-rank subspace and clusters their normalized directions to construct gradient prototypes. It then learns a lightweight router that uses the current input activation to select which prototype to apply at inference time. Across seven behavioral steering tasks and two LLM backbones, GRASP improves seven-task macro accuracy over the unmodified Qwen2.5-7B and Llama-3.1-8B models by 13.37 and 20.23 percentage points, respectively. Ablation studies show that two prototypes outperform a shared direction, while routing controls confirm the need for input-dependent selection.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.