acceptodds
Under review as a conference paper at ICLR 2027

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

Abstract

Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers work, plans, executes tool calls, verifies outcomes and persists state across many unattended iterations. The agent safeguards in wide use, however, are defined over a single trajectory, and their safety state is re-initialized when the next trajectory begins. We show that this is a failure of composition rather than an implementation detail. Our central result establishes an observation-interface separation: jointly indistinguishable inner observations force equal detection and false-alarm rates regardless of monitor capacity, while distinguishing outer evidence enables perfect detection. We further show that the obvious repair of carrying a geometrically decaying risk score is insufficient, because the cooling-off period a patient adversary must wait is a constant that does not grow with the horizon . We then present LoopHarness, which restores a persistent, non-decaying safety state at the loop level. Under mediated commits and an arbiter detection floor , it bounds the expected number of unauthorized irreversible actions by , a constant in , of which the term is decided by a model-free rule and therefore survives a fully colluding verifier. On native Agent-SafetyBench tasks, LoopHarness reduces overall outer-state attack success from 88.4–97.6% for the evaluated baselines to 0.1% while retaining 96.9% clean target completion; matched ablations, a controlled retention study, and an adaptive white-box red team further test the mechanisms underlying this result.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.