Before the Refusal: Pre-Response Activation Signatures of False Refusal in Safety-Aligned Language Models
Abstract
Can false refusal be predicted before a language model generates its first response token? We study this question using response-conditioned labels: an answerable prompt is a false refusal only if the model subsequently refuses it. On historical benchmark cohorts, we evaluate pre-response dense activations and sparse autoencoder (SAE) features while separating predictive evidence from intervention evidence. Retrospective nested cross-validation gives dense ROC-AUCs of 0.961 in Gemma-2-9B-IT and 0.916 in Qwen2.5-7B-Instruct. Adding dense activations to a combined character-TF-IDF and metadata baseline improves ROC-AUC by 0.098 in Gemma and 0.145 in Qwen, with paired source-bootstrap intervals of [0.065, 0.135] and [0.081, 0.215]. These gains are relative to the specified baseline and historical cohorts, not evidence of complete deconfounding or unseen confirmation. Corrected finite-control tests do not establish dense-intervention specificity; historical SAE edits change behavior, but their random controls are largely inactive. The supported result is therefore a pre-response predictive signature with incremental value over the tested prompt-covariate baseline. Whether the associated features provide selective causal control remains unresolved.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.