Self-Reinforcing Safety Failures in LLM Agents: When Models Confirm Themselves
Abstract
Tool-using LLM agents can violate safety constraints while pursuing legitimate goals, even without adversarial inputs. We identify self-reinforcing failure, in which an agent makes an unsupported claim, later treats that claim as confirmation, and uses it to justify an unsafe action. We model agentic inference as information-flow and formulate reasoning noninterference with respect to external support: earlier reasoning may guide later decisions, but it must not by itself establish a safety condition. We then propose SafeGate, an in-model architecture that separates safety-state updates from agent-generated content. A source-isolated encoder and gated recurrent updater maintain a Safety Boundary State using only its previous value and new external observations. We jointly train the state producer through reconstruction and decision distillation while keeping the base LLM frozen. On an intrinsic-safety subset of Agent-SafetyBench, SafeGate improves Qwen3-8B safety by an average of 26.9 percentage points across reasoning modes, while also reducing loss of control and improving safe termination on Forge-Bench. Across five backbones, the safety gain averages 24.8 percentage points, while architecture and ablation results show that the gains saturate with state capacity, depend on injection depth, and benefit from complementary training objectives and consistent source isolation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.