acceptodds
Under review as a conference paper at ICLR 2027

TESTING WHETHER A JAILBREAK DEFENSE REDUCES ATTAINABLE RISK

Abstract

Jailbreak defenses are usually evaluated by giving an attack a fixed number of attempts and measuring how often it succeeds. A lower success rate can mean two different things: the defense made some behaviors unreachable, or it only made them take more attempts. A fixed budget cannot tell these apart, so delay can be reported as prevention. We introduce a certificate that separates the two: it replays queries that already succeeded on the undefended model, and every one that still succeeds proves a behavior was not prevented. We validate the certificate on ANCHOR, a benchmark we construct in which each defense’s true effect is known. A standard evaluation reports a large gain for a reversible input recoding that cannot prevent any behavior, whereas our certificate shows it could have prevented at most 3.6% of them. Stochastic filters leave behaviors reachable but increase the attempts required. The certificate also gives informative bounds for four published defenses. Jailbreak evaluations should therefore report prevention and attacker cost separately.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.