Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential
Abstract
Mechanistic interpretability has emerged as the primary means to understand safety behaviour of LLMs. However, existing tools primarily focus on the activating neurons or features of a model. The role of the remaining large set of inactive components is invisible to such methods. This work demonstrates that the inactive set actually contains safety-critical features that are causally relevant for refusal of harmful prompts. Suppressing such features could turn refusals into compliance, while still passing undetected by prevalent interpretability tools. We introduce the Counterfactual Activation Potential (CAP), a novel metric that quantifies a suppressed feature's latent activation tendency as the product of its encoder alignment (how strongly the input drives it), suppression strength (how strongly active features inhibit it), and safety criticality (how much refusal depends on it), so that the score is high only when all three hold. For efficient discovery of suppressed safety features at scale, we propose CAP-guided Safety Feature Discovery (CSFD), a two-stage filtering algorithm that identifies candidate safety features from hundreds of thousands of transcoder features without exhaustive ablation. In our experiments, a significant fraction of trials turn compliant with harmful prompts when a candidate feature is ablated. We observe that under natural jailbreak conditions, the suppression acting on the highest-CAP features rises 2–4×, and their activation correspondingly falls by up to 80%. Intervening and amplifying a feature's suppressors correspondingly pushes its activation down and raises harmful compliance with prompts related to the suppressed feature, with no such effect for random features. Our experimental validation spans five Gemma, Qwen, and Llama models across both base and instruction-tuned variants. *Our findings indicate that jailbreaks could operate in part by suppressing safety-critical features rather than solely activating harmful ones, and that suppressed features are a necessary complement to activation-focused interpretability of safety behavior.*
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.