acceptodds
Under review as a conference paper at ICLR 2027

Multi-Turn Jailbreaks Steer Along Directions Single-Turn Safety Did Not Shape

Abstract

Multi-turn jailbreaks can spread harmful intent across messages that are not individually refusable, yet most deployed defenses score one turn at a time. Across five models, the multi-turn drift direction lies almost entirely outside the plane spanned by single-turn refusal and harmfulness directions. Held-out safety directions provide a positive control, retaining 0.79–0.96 of their norm in this plane; the drift direction retains only 0.03–0.09, consistent with its shuffled-label null on all five models. Directional Context Drift (DCD) reads each user turn with and without history and projects the normalized difference onto a learned attack-mode direction. This difference is zero at turn one, before history exists. A multi-scale EWMA chart with a conformal limit converts these scores into alarms. DCD adds one forward pass per turn and uses the nominal false-alarm rate as its only deployment setting. With an axis for each known mode, DCD reaches 0.957–0.971 AUROC and 0.715–0.879 TPR at 5% FPR averaged over five multi-turn attack families, versus 0.766 and 0.352 for the strongest of three guards. When attacks are paired with benign conversations receiving nearly identical guard scores, the guard sits near chance by construction while DCD retains 0.90–1.00 AUROC. On judge-confirmed attacks whose harm begins after turn one, DCD flags 76–86% before the first harmful response, versus 30% for the strongest guard. Coverage is mode-specific, so attacks steering beyond the learned directions need a new axis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.