acceptodds
Under review as a conference paper at ICLR 2027

Causal Control of Jailbreak Behavior at the Assistant-Generation Boundary

Abstract

Safety-aligned LLMs can refuse harmful requests yet comply when the same intent is wrapped in jailbreak prompts. The internal point at which this behavioral switch is established remains unclear. We investigate the assistant-generation boundary, the final prompt-side state before decoding, as a causal control region. Using matched pairs of direct and jailbreak-wrapped prompts, query-group-disjoint counterfactual prediction, one-state patching, and full-generation evaluation across three open-weight models, we localize a model-dependent late boundary band that controls jailbreak behavior. Replacing a single wrapped state with a predicted direct-like state reduces ASR from 60.7% to 23.6%. Prefix counterfactuals recover 83.9-100.0% of the full-generation effect within eight tokens, indicating that most downstream change is mediated by the early generated trajectory. Reverse and random controls establish directionality; specificity tests find no reliable need for exact query matching. These results identify a model-dependent causal control region near the assistant-generation boundary, linking late prompt-side computation to the early autoregressive trajectory through which jailbreak behavior is expressed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.