Certifying LLM Safety Across Adaptive Multi-Turn Conversations
Abstract
A large class of language-model safety certificates is local: it bounds the effect of perturbing one prompt. A multi-turn jailbreak is different. The attacker observes each response, adapts the next message, and succeeds if any turn drives the conversation into an unsafe state. We formulate certification for this sequential event and introduce Multi-Turn Certified Robustness (MTCR), a composition layer that lifts valid one-step probability bounds to adaptive conversations. MTCR partitions safe dialogue histories into abstract states and certifies the safe probability mass transferred between them. A robust dynamic program then propagates this mass over the conversation horizon, retaining reachability and state-dependent risk that a global product discards. We further refine the abstraction with safety-margin bins, recording not only whether a trajectory remains safe but how close it moves toward the decision boundary. The resulting certificate requires neither independence across turns nor a fixed attack trajectory. Across four open-weight and two hosted language models, MTCR improves the certified lower bound over global worst-case multiplication at every evaluated horizon; on LLaMA-2-7B-Chat, the improvement grows from at five turns to at twenty turns. Controlled experiments isolate the source of this gain and recover exact survival probabilities when the abstract transition model is known. These results establish state-aware composition as a practical route from single-turn robustness to certifiable conversational safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.