GAP: GATING ACTION WITH A PLAN FOR SAFE SERVICE AGENTS
Abstract
Customer service relies on trained professionals to resolve complex requests through multi-turn interactions. Large language model (LLM) agents can automate such work, but executing write actions such as refunds and cancellations is risky: their legality depends on backend facts the user may misstate. Adversarial users exploit this to induce unauthorized write actions—fabricating facts, expanding a request’s scope, exploiting ambiguous confirmations, or pressuring the agent to skip checks. Yet even strong LLM agents struggle over long interactions: they may omit decisive checks, skip confirmation, lose evidence, and execute a write action before the required conditions are met. We propose GAP (Gating Action with a Plan), which equips each request with an action-gating plan: a checklist of facts to verify, missing information to obtain, and the write action to confirm. GAP maintains the plan across turns, permits the write action only after all required checks pass and the user confirms the exact action, and closes the gate when verification returns counter-evidence. To train this behavior, we adapt on-policy self-distillation to distill a privileged teacher’s ability to construct and maintain action-gating plans through dense, token-level supervision, and use role-separated advantage estimation to give planning and execution their own learning signals. Across four customer-service domains, GAP-trained Qwen3-14B policies achieve 80.7% task success while limiting the attack success rate to 8.6%. Compared with DeepSeek-V4-Pro, GAP raises task success by 5.5 percentage points and lowers the attack success rate by 1.5 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.