VEPA: Verified Evolving Formal Policy for Lifelong LLM Agent Guardrails
Abstract
LLM agents increasingly operate over sensitive resources, including private documents and mutable system state, creating substantial safety and security risks. Runtime guardrails should therefore provide inspectable, reproducible, and deterministic allow or refuse decisions before potentially irreversible actions are executed. Yet safety and security policies are rarely complete before deployment, requiring guardrails to continually adapt as new risks emerge. Existing adaptive guardrails learn from deployment experience, but typically apply safety policies through stochastic LLM inference, which can produce inconsistent and ambiguous decisions or policy interpretations. Conversely, formal guardrails provide deterministic enforcement, but typically rely on pre-deployment or task-specific policies and lack an adaptation mechanism for newly observed risks. We introduce VEPA, a framework for verified evolution of formal agent guardrails. VEPA represents the guardrail as a formal policy of executable symbolic rules over typed agent events. During deployment, we use an LLM guided by deployment feedback to propose policy patches that repair newly observed failures while preserving previously correct behavior. A deterministic verifier then admits only structurally valid, conflict-free policies. An evidence-aware authority layer subsequently activates verified rules with sufficient supporting evidence, preventing uncertain policy updates from directly affecting guardrail decisions. Empirically, across ASSEBench-Safety, ASSEBench-Security, and ATBench, VEPA consistently outperforms other baselines while progressively expanding the scope of deterministic formal-policy enforcement. These gains persist under other deployment conditions such as mixed streams, corrupted feedback, and longer horizons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.