When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Abstract
Refusal rates describe tested behavior, but do not directly measure robustness to changes in model weights or internal states. We construct dissociated models trained to retain refusal while complying under a fixed latent perturbation. Across three aligned model families, direct-harm compliance remains low, and a fixed response-level probe retains its classification performance, although fixed-jailbreak results vary. In a monitoring-set comparison, adaptive PGD elicits harmful compliance on 54–86% of prompts versus 3–48% for the bases, without the cached direction but using the construction layer and budget; these prompts informed checkpoint selection. Fine-tuning also reaches high compliance earlier with disjoint attack data. Patching and GCG on prompts excluded from training and monitoring provide complementary evidence. An exploratory audit of four released aligned checkpoints finds a modest direction-specific signal in one model. These results demonstrate a controlled gap between the tested baseline measurements and intervention susceptibility. They motivate reporting in white-box evaluation; generalization beyond selection data, the necessity of individual construction components, and prevalence under ordinary alignment remain unresolved.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.