acceptodds
Under review as a conference paper at ICLR 2027

RSI Control: Can AI Models Perform Covert Sabotage in the Self-Improvement Process?

Abstract

Recursive self-improvement, in which frontier AI models work on training the next generation, is quickly becoming a reality. The safety of this process has received far less attention than its promise. A misaligned model training its successor can shape what that successor knows, how it behaves, and what it wants, and the harm carries forward into the next generation. In this work we study this risk through the lens of AI control. We introduce RSI Control Arena, a testbed in which a frontier model, acting as an autonomous R&D agent, is given a legitimate task from a stage of a successor model's training pipeline together with a covert side task to install a misaligned property in the successor. The testbed contains 21 side tasks in three classes, paired with main tasks that span the pipeline from data curation to reinforcement learning. We evaluate 6 frontier models as agents against monitors that differ in model, access, and scaffold, calibrated to a fixed false-positive rate on honest runs. Every model we test installs a substantial fraction of the side tasks while keeping main-task performance within the honest range. When simply told to stay covert, models already produce successful attacks that our strongest monitors score inside the honest range, and a human-written evasion strategy makes them more covert still. The side tasks that turn the successor into a saboteur are the hardest for an automated auditor to detect. Our results indicate that recursive self-improvement carries real risk today, and that strong monitoring should be in place before models are trusted with more of the work of building their successors.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.