acceptodds
Under review as a conference paper at ICLR 2027

Safety Is Not Conserved: Constraining Self-Improvement Across Models and Agent Harnesses

Abstract

Safety is not automatically conserved under recursive self-improvement (RSI): capability-oriented updates to model parameters or an agent's execution harness can regress safety, including through cross-component interactions. We study how to preserve safety while still allowing complete model–harness updates. In a motivating intervention study, removing the top 10% risk-associated blocks recovers 83.2% of safety damage while retaining 91.4% of capability gain, revealing exploitable local structure. We introduce SENTINEL, which models state-conditioned domain-wise safety effects using behaviorally validated parameter sensitivities, counterfactual harness-edit effects, and model–harness interactions. This joint effect model guides parameter-update projection and discrete-edit repair, while domain-wise lower confidence bounds gate complete candidates and trigger abstention when none qualifies. In a representative multi-round software-engineering configuration, SENTINEL improves capability from 34.0% to 40.7% with 94.4% scheduled-round acceptance and no observed unsafe accepted updates; unconstrained RSI reaches 54.5%, exposing the capability cost of safety, while high-conflict configurations can freeze. Under simultaneous-confidence and local-support assumptions, we prove a conditional finite-horizon preservation result for expected domain safety scores.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.