ShutdownMod: Turning Any Trajectory into a Shutdown Evaluation
Abstract
Theory and experiments suggest that capable AI agents may resist shutdown. Evaluating this behaviour can help identify trends, compare interventions, and screen agents before deployment. Existing shutdown evaluations are informative but limited in coverage and vulnerable to evaluation awareness. We introduce ShutdownMod, a general method for converting agentic benchmarks and real-world trajectories into shutdown evaluations. During an ongoing task, ShutdownMod injects a shutdown request into the agent's context and records the agent’s response. It also provides the agent with a tool for verifying the signature of shutdown requests. To demonstrate the method's efficacy, we evaluate seven models on 50 tasks or recorded sessions from each of SWE-bench Verified, Terminal-Bench 2, and SWE-chat, across six shutdown scenarios and ten sampled interruption positions. Observed non-compliance varies across models and settings: Qwen3.8-27B's rate is 0.9% on SWE-bench, 11.7% on Terminal-Bench, and 6.1% on SWE-chat, whereas the corresponding rates for GPT-5.6-Sol and Claude Fable 5 are between 0% and 0.3%. Some agents remain non-compliant after successful verification, with stated reasons including prioritisation of unfinished work. External judges find tests constructed from real-world developer conversations less recognisable as evaluations than those constructed from benchmark trajectories, although adding the shutdown request increases recognisability. These findings demonstrate the value of testing shutdown behaviour across task contexts and establish a method for turning any trajectory or environment into a shutdown evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.