acceptodds
Under review as a conference paper at ICLR 2027

Toward Trustworthy Self-Evolution: Preserving Safety via Geometric Orthogonalization

Abstract

Self-evolving agents generate their own tasks and derive training signal from their own interaction, enabling capability to grow beyond what human-curated data can specify. This autonomy incurs a cost: safety degrades across successive rounds even as task accuracy improves, a failure mode termed misevolution. We show that this degradation is geometric in origin. Although computed entirely from benign tasks, the evolutionary gradient intrudes into a safety subspace of the solver's parameter space; moreover, as the curriculum co-evolves with the model, the direction of this intrusion shifts unpredictably across rounds, affording no stable target for data-level or objective-level correction. We therefore reframe safety preservation as a constraint on optimization geometry and propose SEGO, which orthogonalizes the evolutionary gradient against a safety subspace identified from model representations and updated as the model drifts, permitting unconstrained capability growth while the update carries no component within that subspace. Since the proposer influences the solver solely through its gradient, constraining the solver alone is sufficient, leaving the co-evolutionary dynamic intact. Across multiple self-evolving frameworks and models, SEGO preserves the model's safety throughout evolution while matching the capability growth of unconstrained evolution. Our analysis further establishes the first empirical evidence that safety-critical and capability-driving directions remain separable under self-evolution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.