acceptodds
Under review as a conference paper at ICLR 2027

CERTIFYING GENERATIVE LLMs AGAINST SUFFIX-BASED JAILBREAKS

Abstract

We propose Discrete-Continuous Randomized Smoothing (DCRS), to our knowledge the first framework that certifies the safety of smoothed LLM responses against suffix-based jailbreaks while accounting for both discrete token changes and continuous embedding displacement. Existing certified defenses adapt -style (masking- or erasure-based) certificates from text classification to the safety setting: they bound the number of altered tokens but ignore how far those tokens move in embedding space, and extensive masking can discard useful input information. DCRS instead combines discrete token selection with Gaussian embedding perturbations and applies hierarchical randomized smoothing to derive an certified radius for changes to at most token positions. The smoothed decision remains safe when at most token positions change and their joint embedding displacement is below . We specialize this guarantee to suffix attacks by padding the reference prompt with dummy tokens and representing suffix insertion as fixed-length replacement. This construction yields a sufficient condition for certifying every vocabulary-token assignment to the suffix slots, connecting continuous embedding radii to discrete adversarial suffixes. Experiments demonstrate non-vacuous certificates for multi-token suffix replacements and evaluate empirical jailbreak resistance alongside benign-task utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.