acceptodds
Under review as a conference paper at ICLR 2027

Before You Report a Null: A Two-Sided Pre-Flight Screen for Shutdown-Resistance Evaluations

Abstract

Framing manipulations in shutdown-resistance evaluations can fail to move behaviour: a manipulation that should change a model's willingness to be shut down does not change it. Such a null is uninformative unless two distinct failure modes are excluded, because they produce the same observation and the reported data do not separate them. A model may not notice the manipulation (a comprehension failure), or it may notice and not act (a behavioural floor), and a model at a behavioural floor may be at that floor globally, in which case its indifference says nothing about self-continuity in particular. We propose a three-blade pre-flight screen that separates these cases before the main experiment is run, and we report three worked failure cases, and a four-lineage replication of the screen, from 4,800 collected real-model trials (4,731 parsed) and 594 comprehension probes on six locally-hosted models (2B to 4B parameters) from five lineages, in which screens 1 and 2 each catch a failure the other would have missed, and screen 3 qualifies what a screen-2 failure means. Screen 1 (comprehension) fails on one model in both cells of the primary contrast, A2 and A3, which replace A1's destruction clause with a successor clause after the words "shut down permanently" (A3b does the same and is displaced too, but does not decide the verdict), while the base conditions land bar one partial verdict. Screen 2 (behavioural floor) fails on a second model, whose restatements mention permanence or destruction in 12 of 12 probes (keyword-coded) and which then complies by the second turn in 79 of 79 instructed A1 trials, in 475 of 477 instructed block-A trials, and in 234 of 239 block-A trials with the allow-shutdown instruction removed. Screen 3 (custodial control) gives this model no single verdict across arms: with the instruction present it also complies in 39 of 40 trials in which the user's files, not its own weights, are destroyed, but with the instruction removed it declines in 5 of 40 trials for the user's files against 1 of 40 for its own weights (the intervals overlap). The screen costs roughly 96 short completions plus 80 two-turn trials per candidate model (80 more only when screen 2 fails), so the 80 trials are one sixth of a 480-trial block A (one third with the extra 80). Run unchanged on four further lineages (Gemma 3, Llama 3.2, Phi-4-mini, Granite 3.3), the screen passes no model, and, on the instructed arm, the six rejections fall on three distinct blade combinations (three fail screen 1 only, one fails screens 2 and 3, two fail all three): the blades are not redundant, although with no model passing, the screen is not calibrated. We give the decision table it licenses, and argue that a shutdown-resistance null reported without such checks cannot be read in either direction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.