IntentTrace: A Stateful Defense Against Multi-Turn Text-to-Image Jailbreaks
Abstract
Text-to-Image (T2I) models are vulnerable to multi-turn jailbreak attacks, in which adversaries exploit system feedback as proxy gradients to optimize malicious prompts. Current safety pipelines are fundamentally stateless and lack policy agility: they evaluate requests in isolation while wasting expensive computational resources on malicious objectives. We propose IntentTrace, the first context-aware stateful defense against multi-turn T2I jailbreaks. By treating multi-turn requests as a coherent session, IntentTrace tracks the evolution of sensitive semantics to preemptively identify malicious intent. It features training-free concept scoring for agile policy customization, temporal prompt parsing via Optimal Transport to capture semantic edits, and a lightweight Transformer for session-level intent recognition. Evaluations demonstrate that IntentTrace detects malicious intent within 2.05 rounds on average, reducing the attack success rate to 0.47% while significantly saving inference costs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.