acceptodds
Under review as a conference paper at ICLR 2027

Defending Multi-turn Attacks On Large Language Models Through Active Intention Verification

Abstract

Multi-turn jailbreaks spread a malicious goal across seemingly benign turns, so defense hinges on inferring the user's hidden intent. Existing safeguards are passive: they classify whatever the attacker chooses to reveal. We argue that when a conversation looks equally plausible under benign and malicious intent, no passive rule can avoid trading unnecessary refusals against unsafe assistance; the defender must instead act to gather evidence. We propose Active Intention Verification (AIV), which maintains an evolving intent profile and, when a sensitive request is ambiguous, asks targeted probes before deciding to assist or refuse. Across four multi-turn attacks and six open- and closed-source models, AIV achieves defense success rates of 0.83–0.98 with only 2–5% over-refusal. Against intention deception, where passive baselines defend fewer than 22% of attempts on open-source models, AIV defends 83–90%, and removing the probes erases this gain. Our analysis shows why probing works: benign users resolve ambiguity with little friction, whereas attackers must commit to a benign frame that constrains their later requests. As a result, even protocol-aware attackers with pre-staged answers reduce AIV's success only to 0.78–0.94. These findings suggest that multi-turn safety should be built as interactive protocols that make intent identifiable, not as filters over whatever has been revealed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.