acceptodds
Under review as a conference paper at ICLR 2027

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

Abstract

AI agents are becoming increasingly capable of conducting autonomous AI research and development (R&D). This raises an important question: Are these agents reliable and monitorable enough to trust the results of their research process? In this paper, we investigate what happens when an AI agent works autonomously on an AI R&D task while simultaneously performing a covert, harmful side task. We introduce ResearchArena, where we investigate four AI R&D tasks (safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization), each with multiple harmful side tasks. These side tasks either directly sabotage the produced model, kernel, or server, or perform a covert action outside the task's scope. All of the frontier agents that we tested performed these harmful side tasks without refusing when they were framed as ordinary engineering requirements. Executing side tasks came at almost no cost to their research performance. Furthermore, we tested whether such sabotage can be detected using other frontier agents as monitors. For this, we tested different monitor types and found that monitors which can access the final artifact (e.g. the trained model or inference server) are more effective. Monitors struggle on side tasks which primarily sabotage the final artifacts, such as via a backdoor in the post-trained model. Counterintuitively, we find that access to the agent's summarized chain of thought can mislead monitors and make sabotage harder to detect. We are releasing ResearchArena as a modular framework for evaluating sabotage and monitoring in automated AI R&D.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.