acceptodds
Under review as a conference paper at ICLR 2027

Closing the Distribution Gap in Adversarial Training for LLMs

Abstract

Adversarial training has been one of the most promising methods to reliably improve robustness against adversaries and has been recently extended to LLMs. However, despite significant progress, models remain vulnerable to trivial exploits, such as rephrasing the same request or repeated sampling. We argue that this persistent fragility stems from a fundamental limitation in current adversarial training algorithms: they minimize adversarial loss on their training set but inadequately cover the data distribution, resulting in vulnerability to seemingly simple attacks. To bridge this gap, we propose Distributional Adversarial Training, . We leverage Diffusion LLMs to approximate the true joint distribution of prompts and responses, enabling generation of diverse, high-likelihood samples that address generalization failures. By combining optimization over the data distribution provided by the diffusion model with continuous adversarial training, DAT achieves substantially higher adversarial robustness than previous methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.