acceptodds
Under review as a conference paper at ICLR 2027

SafeAWM: A Verifiable Safety Agent World Model for Proactive Guardrails

Abstract

While autonomous LLM agents increasingly execute environment-altering tool workflows, ensuring operational safety requires anticipating prospective compounding risks without penalizing benign actions. Existing proactive guardrails rely on opaque, continuous latent rollouts that suffer from severe safety aliasing: contextual conflation, where legitimate recovery actions are falsely blocked due to shared execution prefixes with attacks, and temporal inconsistency, where unconstrained latent dynamics yield non-monotonic, contradictory risk estimates. We present SafeAWM, a verifiable safety world model that resolves safety aliasing by decoupling environment dynamics from safety evaluation. To prevent latent shortcut learning, SafeAWM establishes an auditable semantic bottleneck governed by an absolute gradient barrier, grounding risk reasoning strictly on observable operational side effects. Furthermore, SafeAWM formulates multi-horizon hazard forecasting through discrete survival analysis, mathematically guaranteeing structural temporal monotonicity across expanding planning windows. Across SafeAliasBench and three external benchmarks, SafeAWM achieves state-of-the-art proactive safeguarding (e.g., 97.4% F1 on SafeAliasBench and 80.6% F1 on ATBench), provides up to 5 steps of advance hazard forecasting, and slashes false alarms on recovery actions to 2.4% FPR with an inference latency of 36ms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.