Rethinking Refusal Directions: Identification and Intervention Effects
Abstract
Low-dimensional activation directions can causally control refusal in large language models, but a direction’s identity and behavioral meaning are often treated as properties of the vector alone. We show that neither is fixed. First, the direction identified as controlling refusal depends on how candidates are constructed and evaluated: lexical and activation-space objectives can select closely aligned or substantially different directions. A validated direction is therefore better understood as an operational causal approximation than as a uniquely recovered internal axis. Second, full-generation semantic judgments analyzed with Signal Detection Theory show that harmful and benign refusal rates change differently across operators, models, and prompt populations, producing changes in both response criterion and the descriptive probit-scale separation index . Third, in a separate continuous-score analysis of contrastive activation addition, we identify a stable prompt-specific downstream gain . Item-specific gains reconstruct local proxy-discrimination trajectories more accurately than class-level or global response models, demonstrating that a common activation displacement can be transformed into heterogeneous functional effects across prompts. Together, these results show that a refusal direction functions as an operational causal interface whose behavioral meaning emerges jointly from its identification procedure, intervention operator, input state, downstream computation, and evaluation population.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.