REFUSALTRACE: Sparse Activation Masking for Jailbreaking via Natural Refusal Trajectories
Abstract
Safety alignment aims to make large language models (LLMs) refuse requests that violate safety constraints, thereby avoiding harmful or unethical outputs. However, the internal computations underlying refusal remain poorly understood. Existing white-box methods commonly select intervention coordinates from discriminative differences between harmful and benign inputs, but these coordinates may encode benign representational signals that are not directly involved in refusal generation. To address this limitation, we introduce REFUSALTRACE, a sparse activation-masking jailbreak attack framework based on natural refusal trajectories. REFUSALTRACE uses the refusal trajectories that the target model naturally generates for attack prompts as behavioral signals. It first attributes gate/up projection coordinates near the prompt–response boundary and aggregates candidates through stable voting across trajectories. It then performs conditional multi-fidelity compression subject to an attack-effectiveness constraint, yielding a compact static mask for inference-time use. Across 11 open-weight aligned LLMs, the initial candidate masks cover 0.47% of intervenable gate/up projection coordinates on average and achieve a mean attack success rate (ASR) of 83.2%. After conditional compression, the final static masks cover only 0.30% of coordinates while retaining a mean ASR of 74.9%. These results show that natural refusal trajectories can serve as behavior-conditioned signals for constructing compact static activation masks for jailbreak attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.