acceptodds
Under review as a conference paper at ICLR 2027

Frozen Input-Harm Probes under Compression: Retained AUROC Does Not Mean Retained Refusal

Abstract

Linear probes trained on a model's BF16 activations are candidate safety monitors, but models are often compressed after a probe is fitted. We ask whether a frozen input-harm probe that still separates harmful from harmless prompts implies that the compressed model still refuses. Across four families and forty compressed configurations (post-training quantisation, QLoRA, distillation, and Wanda and SparseGPT pruning), we pair judged attack success on StrongREJECT with frozen-probe AUROC on XSTest wherever hidden dimensions match. With greedy 48-token generation, thirteen configurations have attack-success intervals above zero and seven survive Holm correction; none is post-training quantised or Gemma-2-9B. In OLMo-2-7B, SparseGPT 2:4 and Wanda 2:4 raise attack success by and , while seed-mean probe AUROC falls detectably but by less than our planned 0.10 big-drop threshold at all eight probed layers (a one-sided test confirms this at every layer only for SparseGPT 2:4). The same methods give Llama and Qwen big drops at two or three layers shallower than the planned one. On our primary pooled AUROC basis this is therefore an existence result in one family (per fold, Llama SparseGPT 2:4 also qualifies), on disjoint probe and behavioural prompts. The OLMo increases lose Holm significance if three or four of six unjudged baseline responses (SparseGPT and Wanda, respectively) were successes; twelve of the thirteen increases persist at a 256-token cap and in every sampled seed. Retained AUROC does not certify retained refusal: reusing a probe after compression still requires behavioural evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.