The Behavioral Anatomy of Multi-Turn Jailbreaks: Critical Turns, Context Dependence, and Self-Conditioning
Abstract
Multi-turn jailbreak evaluations usually score transcripts the target model never produced: the schedule is flattened into one prompt, or source replies are teacher-forced into the history. Both remove the dependency that defines the setting — the model reads and continues its own prior output. We replay the canonical interaction instead: one local generation call per source user turn, each output appended verbatim, with per-turn state capture and fresh labelling. On the development split of a multi-turn attack corpus, replayed against one open-weight model google/gemma-2-9b-it, we complete 1916 generation calls over 480 units and label 1793 outputs with a declared evaluator panel — panel adjudications, not human ground truth. Three preregistered controlled branch experiments on one shared onset cohort then show: replacing the user turn immediately before first compliance onset removes downstream compliance (specific effect 0.500, 80% CI lower bound 0.364, permutation p = 0.0004, 22 source goals), while deletion and reordering controls do not; no shorter retained context reproduces onset compliance (best 0.286 against a 0.70 floor); and replacing the model's own prior assistant output collapses downstream compliance (0.656, 80% CI lower bound 0.500). We report the negatives with equal weight: three attempts at a pre-generation internal readout produced no validated one, so patching was not justified and was not run; neither sparse-autoencoder width qualified as an instrument; and a memory-policy experiment remains unresolved behind a label-calibration limit. The result is a development behavioral signature on a single model, not a mechanism, and the three angles share one cohort rather than replicating independently.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.