ACE-Guard: Automated Co-Evolutionary Guardrails for LLM Security via Inference-Time Natural Language Gradients
Abstract
Prompt injection attacks pose a persistent threat to Large Language Model (LLM) deployments by embedding malicious instructions that override system controls. Existing defenses rely on static rules or require costly model fine-tuning, leaving them exposed when adversarial strategies shift over time. We present ACE-Guard, an automated co-evolutionary security framework that adapts defensive policies entirely without modifying model parameters (). ACE-Guard operates through a dual-phase architecture: an offline co-evolutionary calibration phase that models security adaptation as a repeated game among an Adversary LLM exploring nine structural injection strategies, a Defense LLM equipped with a deterministic Multi-Signal Prompt Inspector (MSPI), and an Evaluation LLM providing bidirectional natural language gradients, followed by an online pre-execution serving phase that enforces safety in a single forward pass. Across 20 co-evolutionary calibration rounds, ACE-Guard achieves an average Defense Success Rate of 83.4% and a Benign Acceptance Rate of 91.9%, yielding a Balanced Harmonic Score () of 87.4%. Extensive evaluations across out-of-distribution attack benchmarks confirm that natural language prompt steering maintains long-term stability, resists adversarial mode collapse, and achieves robust runtime protection without retraining overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.