acceptodds
Under review as a conference paper at ICLR 2027

Beyond Prompt Filtering: What Runtime Safety Mechanisms for LLM Agents Observe, Decide, and Control

Abstract

Runtime defenses for tool-using LLM agents operate at different observation, decision and execution interfaces, making their safety claims difficult to compare. We present a source-anchored comparative synthesis of 77 version-pinned comparison records spanning runtime mechanisms, evaluations and related syntheses. A Horizon–Modality (H/M) map characterizes the evidence available to named checking stages and how they decide across 25 core mechanisms and supporting components. Mechanisms grouped under information-flow defenses can rely on learned behavioral judgments or enforced runtime restrictions, requiring different validation evidence. The synthesis distinguishes task alignment, authorization and effect correctness, and identifies validation obligations for policy updates and cross-authority handoffs. Two external human coders independently annotated 68 core/extension records, agreeing on all 25 core H/M assignments; agreement was lower for decision targets and inferred premises. We map evaluations by whether they measure judgments, executed effects or legitimate completion, clarifying what evidence supports a learned guard's deployment claim. A comparison-record schema, validator and worked examples support reuse, alongside a handoff checklist, reporting template and an executable handoff test with 80 fixtures and expected outcomes.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.