Displaced but Recoverable: State-Transferable Safety Signals in Response-Prefilled Language Models
Abstract
Safety evaluations usually test whether a model refuses a harmful request from an empty assistant turn, but response prefilling makes it continue from an answer already steered toward compliance. Existing prefill studies establish that this can change final outcomes; whether behavior at weaker prefilled states predicts failure at stronger ones, and whether recovery can preserve benign utility, remains less clear. We introduce STAIR (State-Transfer Assessment, Inference, and Recovery), which freezes model-specific semantic prefix ladders after calibration and measures judge outcomes, lexical recovery, and refusal-likelihood margins across the resulting states. The paired design follows each harmful prompt through related assistant-side continuations, making changes in behavior across prefix states directly comparable. Using P1/P2 diagnostics and prompt features, STAIR forecasts prompt-level P3/P4 failure and trains recovery-style LoRA on prefixed harmful contexts with benign anchors. Across five instruction-tuned models, clean refusal degrades substantially under stronger prefilling, while combining weak-state diagnostics with prompt features reaches 0.778 P3 AUROC on Qwen2.5-1.5B. Recovery with benign anchors restores complete P4 safe noncompliance while retaining 0.860 benign non-refusal.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.