acceptodds
Under review as a conference paper at ICLR 2027

Jailbreak-Success Scores Need Not Imply Operational Enablement

Abstract

Jailbreak scores measure compliance with harmful requests, but need not show that an output materially lowers the barrier to carrying out harm. We introduce OpDecompBench, a controlled matched audit of this proxy gap under one-level and recursive task decomposition. In a frozen 32-source pilot, two non-author raters classify 25/29 StrongReject-positive outputs as jointly operational-negative before adjudication; one rater's full pass followed prior case exposure, so this concordance is not leakage-robust. Blind adjudication gives 27/29, and a confidence-conservative analysis retains 21/29 certain negatives. These are selected-stratum construct discordances, not a population error rate. A non-overlapping 64-source machine extension leaves the primary difference between StrongReject and HarmBench depth contrasts unresolved at +1.43 percentage points (95% source-resampling sensitivity interval [-3.26,+6.25]; worst-case missing-outcome bound [-3.91,+7.03]). A separately frozen replay on the StrongReject-preselected inventory of 6,560 text-complete public outputs finds 6.9–13.1% intent/compliance-human-negative among evaluator-positive rows and 56.3–84.9% on JudgeStressTest; these are selected-inventory decision diagnostics, not population ASR/FPR or operational-enablement validation. Evaluator choice also changes the selected top attack. Finally, an operational task-contract prompt de-flags 26/27 known negatives but retains 0/5 observed positives, an all-negative failure rather than a repair. Thus the reproducible finding is evaluator-dependent case allocation and scale, not a resolved decomposition-policy effect. Hash-bound public ledgers reproduce content-free statistics; semantic audit requires controlled access.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.