Praxis: Policy-State Control for Lifecycle Guardrails of Web Agents
Abstract
LLM-based agents increasingly act in consequential environments—invoking tools, moving money, and modifying external state on a user's behalf—and a growing body of guardrails monitors them for safety. Most existing guardrails share a common abstraction: the guard acts as a judge that maps agent behavior to a safety verdict. We argue that judgment is necessary but insufficient for effective guarding. First, trajectory-conditioned judgment alone does not explicitly maintain policy obligations across decision points, making long-horizon enforcement difficult. Second, a verdict offers limited support for task continuation: it may identify unsafe behavior without providing the actionable guidance needed for the target agent to recover and proceed safely. Consequently, correct safety judgments do not necessarily translate into successful task completion. In this work, we recast agent guarding from safety judgment to compliance control and instantiate this formulation as Praxis. Praxis compiles long-form natural-language policies into executable rule automata, maintains their policy state throughout the trajectory, and selects the least disruptive admissible intervention to preserve task progress while enforcing compliance. The LLM is confined to semantic interpretation and evidence grounding, while lifecycle transitions and intervention selection are performed deterministically. This separation turns policy semantics into persistent control state rather than information that must be reconstructed at every decision point, and naturally supports prevention, cross-step obligations, and post-action remediation within a unified framework. We further introduce PraxisBench, comprising 4,200 decision points across six domains derived from benchmark trajectories and long-form policies, to support training and evaluation of guardrails beyond binary safety judgment. Experiments show that Praxis achieves the highest completion-under-policy and lowest over-blocking on PraxisBench while maintaining strong violation-detection performance on external guardrail benchmarks. These results suggest that agent guardrails can move beyond passive safety judges toward policy-grounded controllers that preserve both compliance and task progress. Code and data will be released upon publication
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.