Wide Exploration, Localized Escape: High-Dimensional Branching Dynamics of Multi-Turn Jailbreaks
Abstract
Safety alignment mechanisms for large language models (LLMs) are commonly optimized against isolated or single-turn malicious requests, while adversaries can distribute harmful intent across multiple adaptive interactions. We study multi-turn jailbreaks from a spatial-dynamical perspective and formulate their evolution as a high-dimensional branching random walk with adversarial drift. In this framework, adaptive prompting steers semantic trajectories toward a proxy safety boundary, generation uncertainty introduces stochastic perturbations, and parallel conversational paths form an expanding attack tree. The model predicts that branching amplifies the finite-sample semantic frontier and accelerates first boundary crossing, while successful crossings concentrate around a limited set of localized exits. We evaluate these predictions using two multi-turn jailbreak methods across multiple mainstream LLMs and semantic encoders. Increasing the number of parallel paths consistently expands the reachable semantic frontier, advances the first crossing, and raises the probability of observing at least one crossing within a fixed interaction budget. These gains arise primarily from repeated path exploration rather than greater semantic diversity among branches. Moreover, successful crossing points remain closer to localized exit regions than matched non-crossing controls, and their terminal transitions exhibit stronger directional alignment with the corresponding exits. Although exact exit locations depend on the semantic representation, these localization and alignment patterns remain stable across encoders. Finally, ablations show that combining exit directions, local semantics, successful examples, and regional context yields more reliable attack-reconstruction gains than any isolated signal. Our findings reveal a characteristic structure of multi-turn jailbreaks—wide exploration inside the safety region followed by localized boundary escape—and provide a mechanistic account of their structural risk.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.