acceptodds
Under review as a conference paper at ICLR 2027

From Safety Memory to Enforced Policy: Continuity and Scope Across Agent Replicas

Abstract

Adaptive agent defenses learn safety rules from past interactions to prevent recurring harm. Yet remembering a rule does not ensure that later actions obey it. We study this gap in serverless agents, where learning and execution can occur on different replicas. We propose Service-Global Monotonic Immunity (SGMI), which shares accepted rules across replicas, defines their action scope, and blocks matching tool calls before generative reconsideration. In Knative scale-out experiments, shared memory retrieves the learned rule on every fresh-replica probe but still permits 47.9% of harmful actions. SGMI blocks all 144 tested harmful probes across original and fresh replicas. In a paired banking ablation, recipient boundaries reduce legitimate-action rejection from 62.5% to 25%, while preserving all eight harmful-action blocks. SGMI connects adaptive safety learning with persistent, action-level protection across changing execution instances.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.