Conf-Guard: Dynamic Multi-Turn Jailbreak Detection with Conformal Guarantees
Abstract
As the applications of large language models (LLMs) expand to critical domains, their capacity for abstention becomes increasingly vital. Current methods treat abstention mostly as a single-turn problem, rather than a conversational one. This issue is of particular importance for securing LLMs against adversarial threats, which have recently shifted from single-turn attacks to sophisticated conversational exploits where malicious intent gradually unfolds across multiple turns. Traditional defences struggle to guard against multi-turn attacks and do not offer formal guarantees on the percentage of attacks they block. Here, we propose Conf-Guard, a novel framework that operates by keeping track of a scalar score which dynamically evolves throughout the conversation, and by terminating the conversation if the score breaches a threshold. We show that the threshold can be selected so as to achieve a distribution-free conformal guarantee, which, for jailbreak detection, formally ensures a user-specified percentage of attacks is successfully blocked. Extensive empirical evaluations demonstrate that Conf-Guard consistently outperforms leading jailbreak detection baselines while satisfying its conformal guarantee. Warning: the data we use in this paper contains jailbreak attacks with toxic and offensive language.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.