Positional Inference for Stronger Attacks on Language Models
Abstract
Optimization-based adversarial attacks on large language models (LLMs) perform coordinate descent over a fixed set of token positions. While seminal works focus exclusively on suffix-based attacks, recent studies suggest that inserting tokens into alternate positions (prefix and infix) can achieve higher attack success rates (ASR). Although methods that move beyond suffixes often improve ASR, their position selection usually relies on architecture-dependent heuristics. We cast position selection as Bayesian inference: intermediate optimizer evaluations that existing attacks discard serve as observations about which positions are worth attacking, and existing heuristics can be formulated as prior beliefs about impactful positions. Because the resulting posteriors are Gaussian, beliefs inferred from individual attacks can be merged in closed form, so computation spent attacking one model becomes a positional prior for attacking another. Across 12 LLMs ranging from 0.6B to 32B parameters, and despite the exploration cost of jointly inferring positions and optimizing tokens, our attack improves mean ASR by 11.1 percentage points over suffix-based GCG and 3.6 points over attention-guided SlotGCG. The inferred posteriors reveal that vulnerability concentrates at both ends of the prompt, reconciling conflicting reports that favor suffixes or prefixes, yet correlates only weakly with the attention signal current heuristics rely on. Most interestingly, we find that positional knowledge transfers: in leave-one-model-out experiments, using posteriors merged from the other 11 models as a prior improves a gradient-free attack on the held-out model by 9–11 points over suffix placement and 6–7 points over transferred attention heuristics, matching suffix-only ASR with roughly a quarter of the query budget. Our results suggest that where to attack is learnable, transferable, and enables an important form of adaptive attacks for language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.