acceptodds
Under review as a conference paper at ICLR 2027

REFUSALTRACE: Sparse Activation Masking for Jailbreaking via Natural Refusal Trajectories

Abstract

Safety alignment aims to make large language models (LLMs) refuse requests that violate safety constraints, thereby avoiding harmful or unethical outputs. However, the internal computations underlying refusal remain poorly understood. Existing white-box methods commonly select intervention coordinates from discriminative differences between harmful and benign inputs, but these coordinates may encode benign representational signals that are not directly involved in refusal generation. To address this limitation, we introduce REFUSALTRACE, a sparse activation-masking jailbreak attack framework based on natural refusal trajectories. REFUSALTRACE uses the refusal trajectories that the target model naturally generates for attack prompts as behavioral signals. It first attributes gate/up projection coordinates near the prompt–response boundary and aggregates candidates through stable voting across trajectories. It then performs conditional multi-fidelity compression subject to an attack-effectiveness constraint, yielding a compact static mask for inference-time use. Across 11 open-weight aligned LLMs, the initial candidate masks cover 0.47% of intervenable gate/up projection coordinates on average and achieve a mean attack success rate (ASR) of 83.2%. After conditional compression, the final static masks cover only 0.30% of coordinates while retaining a mean ASR of 74.9%. These results show that natural refusal trajectories can serve as behavior-conditioned signals for constructing compact static activation masks for jailbreak attacks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.