Refusal Is Not Failure: From Model Responses to Safe System Traces
Abstract
A correct refusal can still lead to an unsafe action when a routed language-model system treats it as a failed call and retries another backend. We model generation, routing, and external actions as a single execution process. This gives an exact expression for system risk, proves that retry after safe termination adds nonnegative unsafe probability, and shows why per-model safety scores cannot determine chain risk. Across four instruction-tuned model families and all 650 XSTest and JailbreakBench prompts, naive retry recovers 394 safe answers but adds 61 responses labeled unsafe by WildGuard and 161 by Llama Guard over 2,600 paired units. Typed-action experiments reproduce refusal bypass with and without token constraints. Under unrestricted decoding, checking each action against the original policy before execution recovers 589 additional safe completions over exact fail closed, with zero unsafe commits. With peer messages, these checks retain 95.4% of naive retry's correct completions while preventing all 95 unauthorized commits on the same finite policy grid. Checking each reached commit against the original policy enables safe continuation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.