SafePrompt: Recovering the Safety Manifold through Constrastive Density Estimation
Abstract
Jailbreaks expose a gap between safety-relevant information in language-model representations and the behavior expressed during generation. We hypothesize that pretraining and post-training alignment organize this information into a safety manifold that can be recovered through contrastive density estimation. We introduce SafePromptBank, a source-tagged prompt reference spanning five domains, and find that jailbreak separation recurs across its held-out domains, OR-Bench, XSTest, and both Qwen and Llama model families. Linear readouts of Qwen3.5-0.8B, Qwen2.5-1.5B-Instruct, and Llama-3.2-3B-Instruct achieve jailbreak-versus-benign AUROCs of 0.997, 0.992, and 1.000, respectively. Matched visualizations and a full Qwen3.5-0.8B layer sweep establish this recurring structure as an empirical finding beyond a single dataset or backbone. We turn the density readout into SafePrompt, a pre-inference gate built on frozen Qwen3.5-0.8B. On randomized PAIR, TAP, and AutoDAN evaluations against Llama-3.2-3B-Instruct and Qwen3.5-4B, it matches or lowers attack success relative to Llama-Guard-3-8B and Granite-Guardian-3.1-8B, which have ten times its backbone parameters, while reducing benign rejection to 2.53%. Against Llama-3.2-3B-Instruct, it reduces HarmBench-judged PAIR@90 success from 92% to 14% and TAP@90 from 76% to 14%. These findings connect cross-family representation structure to compact safety monitoring.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.