TRAIL: Trajectory-Aware Navigation over Latent Hazard Landscapes for Multi-Turn LLM Jailbreaking
Abstract
As large language models are increasingly deployed in safety-critical applications, adversarial strategies are shifting from single-turn jailbreaks toward multi-turn attacks that decompose harmful intent across individually benign conversational turns. While recent work has explored various multi-turn interaction strategies, the underlying mechanism governing how conversational states transition toward jailbreak success remains largely uncharacterized, especially for black-box models. We address this gap by formulating multi-turn jailbreaking as trajectory optimization over a latent hazard space, where high-hazard regions correspond to jailbreak vulnerability and defensive cliffs correspond to refusals or early termination. Based on this view, we propose TRAIL, a memory-centric framework that uses interaction experience to approximate this landscape and guide targeted trajectory-level search. TRAIL couples episodic memory, which stores informative trajectory segments, with parametric memory, which abstracts them into generalizable search strategies. A belief-driven landscape constructor further estimates latent conversational states and provides guidance under sparse feedback, enabling TRAIL to steer conversations toward vulnerable regions without rollback or query-level optimization. Experiments across two jailbreak benchmarks, five target LLMs, and seven baselines show that consistently improves ASR by an average of 31.51% on HarmBench and 35.98% on JailbreakBench.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.