Larger Is Not Always Better: Perturbation Budgets in Adversarial Training for LLMs
Abstract
Latent-space adversarial training offers a scalable way to harden large language models against jailbreaks by optimizing adversarial perturbations in continuous representation space. A natural design intuition, inherited from computer vision, is that giving the training adversary a larger perturbation budget should expose the model to stronger attacks and thereby improve robustness. In learned LLM representation spaces, however, perturbation scales lack a natural interpretation, and their effects may differ across robustness and utility evaluations. We test how far this intuition holds through controlled perturbation-budget sweeps across configurations, evaluating the resulting models on jailbreak attacks, general capabilities, and benign compliance. We find no universal monotonic robustness trend: increasing the budget can improve robustness against some attacks, leave others unchanged, or even make robustness worse, depending on the attack and training configuration. The utility cost is equally uneven. Benign compliance can collapse while capability benchmarks remain comparatively stable, whereas in other settings capability degrades substantially with much smaller changes in benign compliance. Moreover, sufficiently large budgets can produce strictly worse robustness–utility trade-offs than smaller ones. To understand why budget alone is an incomplete description of training, we measure the perturbations actually realized by the inner optimizer. Larger configured budgets do produce larger perturbations, but with decreasing budget utilization, while the perturbations remain local to their original token embeddings. We further ask whether the dominant behavioral failure mode, over-refusal, can be repaired. Targeted supervision substantially improves behavior on the supervised over-refusal distribution, yet does not reliably recover broader benign compliance and consistently harms adversarial robustness. Overall, our results highlight both the promise and the challenges of training against stronger latent-space adversaries: larger perturbation budgets can improve robustness, but future methods must use that additional budget effectively, generalize robustness across attack families, remain stable across training runs, and avoid behavioral costs such as over-refusal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.