acceptodds
Under review as a conference paper at ICLR 2027

Enabling Certified Robustness for LLMs at Large Perturbation Radii via Corruption-Rate Randomization

Abstract

The widespread deployment of large language models (LLMs) makes jailbreak attacks that induce harmful outputs a critical safety concern. Unlike empirical defenses, randomized smoothing (RS) provides robustness guarantees against all attacks within a specified perturbation budget, but its application to LLMs often certifies only a few token modifications despite extensive sampling. We identify a fundamental bottleneck in widely used RS mechanisms based on fixed-probability masking and deletion: even when every sample receives the target label, the certifiable radius is subject to a ceiling that grows at most logarithmically with the sampling budget, severely limiting certification against large perturbations. To address this bottleneck, our method extends randomization from choosing which tokens to corrupt to choosing the corruption rate itself. By mixing lightly and heavily perturbed inputs while preserving the expected corruption rate, it increases the probability of removing multiple attack-related tokens together. We show that this change can reduce the sampling requirements for certification from exponential to polynomial in the target radius. Experiments on LLM-based harmful-content detection demonstrate substantially larger certified suffix lengths at the same sampling budget, showing that our method extends certified robustness to larger perturbation radii.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.