acceptodds
Under review as a conference paper at ICLR 2027

SafelySample: Safety Steering via Contrastive Soft Prompting

Abstract

Autoregressive language models can produce unsafe completions when harmful requests bypass their alignment safeguards. Existing defenses mitigate this risk, but often require extensive model training, auxiliary scorer or expert models, or substantial losses in performance on benign tasks. We introduce SafelySample, a decoding-time method that steers generation toward safer alternatives using a safety signal obtained from the language model itself. We train a small soft prompt on paired safe and unsafe responses while keeping the underlying model frozen and preserving its behavior on benign instructions. During generation, SafelySample compares the model's default next-token distribution with the distribution induced by the safe prompt. It promotes tokens favored by the safe prompt and suppresses tokens associated with the model's default, potentially unsafe continuation, while remaining anchored to the default distribution. SafelySample therefore requires neither an external safety model nor a separately trained unsafe model. We evaluate eight instruction-tuned models from five families on ten benchmarks spanning safety, reasoning, and helpfulness. SafelySample reduces mean attack success rate by 28-45% on six models while keeping average reasoning and helpfulness consistent. We thus demonstrate that SafelySample can steer a foundation model toward safer generation while largely preserving its general capabilities.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.