Breaking the Refusal Barrier: Stochastic Discovery of Residual Unsafe Behaviors in Aligned Language Models
Abstract
LLM safety alignment often creates refusal modes: dominant, high-probability regions of the output space that steer generation toward refusals, while residual unsafe behaviors persist in low-probability regions. Alignment suppresses harmful responses but need not eliminate these underlying behaviors, raising a practical question for iterative safety hardening: how can residual unsafe behaviors be efficiently discovered and converted into safety-training data? We formulate this as stochastic exploration rather than adversarial prompt optimization, showing that random input perturbations redistribute the model's response distribution away from dominant refusal modes and toward low-probability unsafe modes. Building on this, we introduce SEEK (Stochastic Exploration to Elicit residual-behavior Knowledge), a simple inference-only framework that repeatedly samples perturbed inputs and responses to discover diverse residual unsafe behaviors, without gradient optimization, semantic prompt engineering, or auxiliary attacker LLMs. We cast discovery as a coverage objective and show that coverage increases monotonically with the sampling budget, asymptotically recovering all reachable unsafe behaviors. Across five safety-tuned LLMs and two benchmarks, SEEK's discovery matches or exceeds white-box optimization at larger budgets while using one to two orders of magnitude less compute, and consistently exceeds more efficient optimization-, embedding-, transfer-, and attacker-LLM-based methods. Finally, fine-tuning on SEEK-discovered failures substantially improves robustness, cutting attack success from to at most across optimization-, embedding-, transfer- and attacker-LLM-based attacks, while preserving general task performance. Stochastic exploration thus offers a simple, efficient alternative to adversarial prompt optimization for residual-behavior discovery and safety-data generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.