TESTING WHETHER A JAILBREAK DEFENSE REDUCES ATTAINABLE RISK
Abstract
Jailbreak defenses are usually evaluated by giving an attack a fixed number of attempts and measuring how often it succeeds. A lower success rate can mean two different things: the defense made some behaviors unreachable, or it only made them take more attempts. A fixed budget cannot tell these apart, so delay can be reported as prevention. We introduce a certificate that separates the two: it replays queries that already succeeded on the undefended model, and every one that still succeeds proves a behavior was not prevented. We validate the certificate on ANCHOR, a benchmark we construct in which each defense’s true effect is known. A standard evaluation reports a large gain for a reversible input recoding that cannot prevent any behavior, whereas our certificate shows it could have prevented at most 3.6% of them. Stochastic filters leave behaviors reachable but increase the attempts required. The certificate also gives informative bounds for four published defenses. Jailbreak evaluations should therefore report prevention and attacker cost separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.