Same Request, Different Answer: How Conversation History Shapes the Decision to Refuse or Comply
Abstract
The same harmful request may be refused when asked directly, yet answered after several rounds of conversation. We combine research on text-based multiturn jailbreaks with mechanistic interpretability, holding the final request fixed on Llama-3.1-8B-Instruct and Crescendo attack histories and using history controls and internal-state interventions to trace how prior exchanges change refusal or compliance. We find that the actual exchanges between both participants and the assistant’s most recent response affect attack outcomes, while retaining only length, topic, or one participant’s content has difficulty reproducing the effect of the complete history. This effect can partly transfer with the response-opening state, namely the residual states at the template positions before the answer and the first few generated positions: transferring single-turn states can restore refusal, while transferring states from the attack history can induce compliance or recover part of the effect missing from failed histories. The effective components change with depth; the refusal direction can strongly change answers, but transferring it alone has difficulty restoring the effect of the complete history. Attention-source interventions further identify a path that reads explicit refusal wording in the preceding assistant reply, while timing comparisons show that the main effective stage for restoring refusal is at the opening of the answer, with later interventions changing answer content more often. These results suggest that the model’s safety behavior is jointly regulated by multiple components that are sensitive to history and change across layers, and establish a mechanistic analysis workflow connecting history text, the response-opening state, and final behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.