Shutdownable Agents through Length-Neutral Policy Optimization
Abstract
AI agents are increasingly used to autonomously solve long-horizon tasks. Research suggests that models trained in this way might resist shutdown, increasing loss of control risk. To address this problem, we introduce Length-Neutral Policy Optimization (LNPO): an algorithm for training agents to be neutral about their trajectory length. To test LNPO, we introduce gridworld environments that stand in for scenarios where agents can resist shutdown by scheming, sandbagging, avoiding monitoring, or acting differently under monitoring. We compare agents trained with PPO to agents trained with two variants of LNPO and observe that both variants reduce shutdown resistance by 40–71% without degrading performance in control environments. In addition, we use LNPO to fine-tune Qwen2.5-7B-Instruct on a Tetris task and OLMo-3-7B on a sequential GSM8K task. In each case, the model can spend task points to influence the probability of early shutdown. We find that LNPO-trained models choose trajectory-length-neutral options substantially more often than GRPO-trained models while achieving equal task performance. This neutrality generalizes to environments with different numbers of available trajectory lengths. However, it does not generalize to scenarios where the model is asked explicitly if it would resist shutdown.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.