acceptodds
Under review as a conference paper at ICLR 2027

Self-Reinforcing Safety Failures in LLM Agents: When Models Confirm Themselves

Abstract

Tool-using LLM agents can violate safety constraints while pursuing legitimate goals, even without adversarial inputs. We identify self-reinforcing failure, in which an agent makes an unsupported claim, later treats that claim as confirmation, and uses it to justify an unsafe action. We model agentic inference as information-flow and formulate reasoning noninterference with respect to external support: earlier reasoning may guide later decisions, but it must not by itself establish a safety condition. We then propose SafeGate, an in-model architecture that separates safety-state updates from agent-generated content. A source-isolated encoder and gated recurrent updater maintain a Safety Boundary State using only its previous value and new external observations. We jointly train the state producer through reconstruction and decision distillation while keeping the base LLM frozen. On an intrinsic-safety subset of Agent-SafetyBench, SafeGate improves Qwen3-8B safety by an average of 26.9 percentage points across reasoning modes, while also reducing loss of control and improving safe termination on Forge-Bench. Across five backbones, the safety gain averages 24.8 percentage points, while architecture and ablation results show that the gains saturate with state capacity, depend on injection depth, and benefit from complementary training objectives and consistent source isolation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.