Escaping Alignment: A Physical Trap Model of Best-of-N Jailbreaking
Abstract
Best-of- jailbreaking (BoN) bypasses safeguards of aligned models by drawing independent augmentations of an unsafe prompt and sampling completions of each. Previous works have shown that the attack success rate (ASR) seems to follow a power-law in , which we challenge. The exponent drifts with , with an exponential crossover which is a finite-size artifact of the adversarial dataset. Little work has been done to explore the entire two-budget () attack surface as well as its dependence on the generation temperature . We introduce a simple barrier model where each prompt has a baseline safety level and each augmentation a random thermally activated barrier. Then four numbers, each backed by an interpretable safety mechanism, determine the entire () attack surface. They extrapolate predictions from to , collapse five distinct models on the same scaling function and predict ASR at different temperatures from the one they were fitted at.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.