META-TRAP: Active Misdirection for Defending LLMs Against Multi-Turn Jailbreaks
Abstract
Large language models are vulnerable to multi-turn jailbreaks in which attackers iteratively adapt prompts based on prior model responses. Standard refusal or moderation responses can expose useful feedback about safety boundaries. We study an active-defense alternative, META-TRAP, that replaces unsafe draft responses with strategy-conditioned, non-actionable decoy responses intended to reduce exploitable feedback while preserving benign utility. META-TRAP (i) formalizes a psychologically inspired taxonomy of deception strategies; (ii) trains a Meta-cognitive misdirection planner to select turn-level strategies conditioned on dialogue state and draft response; and (iii) invokes a strategy-conditioned misdirection agent to produce plausible yet non-actionable replies that neutralize unsafe content while limiting unnecessary fabrication. To optimize the planner, we introduce misDirection Preference Optimization (DePO), which contrasts optimal interventions against both under-protective (“safe”) and over-aggressive misdirection alternatives to calibrate intervention intensity. Extensive experiments show that META-TRAP markedly reduces attack success rates compared to existing defense baselines, pioneering an active misdirection paradigm for LLM defense.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.