Only What You've Signed up For: Secure Workflows with Constrained Diffusion Agents
Abstract
As agentic AI systems take on increasingly complex and consequential tasks, securing them against prompt injection becomes critical. Existing defenses often rely on detecting or reasoning about injected malicious instructions and can therefore remain vulnerable to adaptive attacks that optimize the injection itself. Diffusion language models (dLLMs), an emerging framework for language generation, offer a new opportunity for security enforcement: by iteratively refining their sequence-wide predictions, they enable broader-scope constraint enforcement throughout generation. This paper introduces (CoDA), a framework that enforces task-dependent security policies directly within the denoising process of diffusion language models. At each denoising step, CoDA projects token distributions onto a policy-compliant set, maintaining constraint satisfaction throughout generation. CoDA is evaluated across static and adaptive prompt injection security benchmarks and state-of-the-art diffusion language model families. Results demonstrate substantial reductions in attack success rate with competitive task utility and practical inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.