acceptodds
Under review as a conference paper at ICLR 2027

Scalable Intent-Aware Safety for Adversarial LLM Dialogue

Abstract

When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is one of the central challenges of LLM safety. Assuming the worst harms legitimate users; always assuming the best can be exploited. The problem is compounded in multi-turn dialogue, where an attacker's true intent may only reveal itself gradually across many exchanges. Yet, existing safety frameworks apply a contextual bandit treatment without considering the trajectory of the conversation. To that end, we propose Dialogue Critic Guided Sampling (DCGS), a framework that addresses this by inferring user intent at every turn of dialogue. DCGS learns what inferences about the user's intent are probable based on conversational history, and resamples responses accordingly. Formally, we model adversarial dialogue as a Markov Decision Process and learn value and regret-based critics at both the individual token and utterance (full response) levels, scoring candidate responses via an action-value critic. We prove that this inference-time reweighting approximates exponential tilting of the base policy, guaranteeing improvement in expected return for any finite candidate pool, a property that group-relative objectives do not exhibit. Evaluated on CARES-18k, WildJailbreak, Redbench, SafeDialBench, and Harmbench, DCGS outperforms strong robust baselines and frontier models on adversarial dialogue tasks. Once trained, DCGS transfers to frontier models, improving their robustness without fine-tuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.