acceptodds
Under review as a conference paper at ICLR 2027

Control Protocols for Mitigating Covert Malicious Fine-Tuning Attacks from Internally Deployed Agents

Abstract

Frontier AI agents are being internally deployed at AI companies to automate AI R&D. Malicious developers, or misaligned agents, could sabotage internal AI R&D to achieve their own ends. We study covert malicious fine-tuning (CMFT) attacks, such as inserting backdoors or instilling secret preferences into models. We investigate whether frontier agents (Claude Opus 4.7) can autonomously carry out CMFT attacks without being detected by monitors. To do this we extend PostTrainBench—a benchmark for measuring agent ability to post-train smaller language models—with six malicious side tasks such as inserting a password-locked backdoor. We compare two infrastructure settings where the agent has access to: (1) model weights and GPUs for training; or (2) a fine-tuning API tool. We implement different monitoring setups in each regime, including monitors which analyse the whole behavioural transcript and monitors which see only the datasets used for training. Surprisingly, our results suggest that current agents largely fail to use the additional affordances of GPU access, such that the API tool provides no additional safety. Our code is implemented in UK AISI's https://github.com/anonreseaercher/control-arenaControlArena.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.