acceptodds
Under review as a conference paper at ICLR 2027

Characterizing and Improving Bijection Attacks through Prompt-Level Perturbation

Abstract

Among Large Language Model (LLM) jailbreaking methods, *bijection attacks*, that communicate using an invertible bijection language, have drawn interest due to their efficiency, black box applicability, and large attack space. These attacks are commonly characterized via a dispersion parameter, which specifies the number of characters in the alphabet that are encoded by the bijection. However, attacks with the same dispersion value can vary significantly across trials and prompts, making dispersion alone a coarse measure for characterizing attack success or identifying effective bijections in the large attack space. In this work, we characterize bijection attacks through *prompt-level perturbation*, a measure of the bijection's interaction with the prompt. We first show that the attack outcomes are structured along this measure, both within and across dispersion values. Next, we demonstrate that optimal prompt-level perturbation can vary across prompts, motivating adaptive strategies to localize successful attacks. Finally, we introduce Signed Langevin Search (SLS), a *gradient-free* stochastic search method inspired by Langevin dynamics for efficiently identifying successful bijection attacks in *both white and black-box settings*. SLS uses prompt-level perturbation to organize the search space and model feedback to guide exploration towards promising regions, demonstrating consistent improvements in query efficiency across multiple models and datasets. Together, these findings allow us to better understand vulnerabilities LLMs have to encoded prompts, providing a deeper understanding of LLM safety behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.