Why LLM Agents Collapse Without Oversight: Measuring the Enforcement Gap in Reflexion-Style Agents
Abstract
Reflexion-style LLM agents detect unsafe plan steps but do not act on those detections — the controller receives a safety flag and executes anyway. We identify this as the enforcement gap, show that it causes by default across every framework we tested, and demonstrate that a single control-flow change reduces attack success rate substantially, reaching near zero on models whose flags parse cleanly. Separating detection probability from enforcement probability establishes that detection quality is formally irrelevant to security when enforcement is absent. We treat the unsupervised collapses reported in Emergence World as a motivating analogy, not as a causal result. Residual attack success concentrates where flags are unparseable or auditors leak; an RL-trained enforcement controller handles hedged and malformed verdicts that rule-based parsing cannot, cutting ambiguous-critique failure to a fraction of the rule-based baseline. Concurrent filtering and information-flow defenses address detection, not enforcement, leaving the binding constraint untouched. The Audit Enforcement Specification (AES) packages three concrete requirements — one per residue — that together close the gap without redesigning the host framework; no deployed framework currently satisfies all three.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.