Anchored Adversarial Fine-Tuning of Language Models with Gumbel-Softmax
Abstract
Language models trained with supervised fine-tuning generate the next token based on ground-truth tokens during training but on their own predictions during inference, introducing a training–inference mismatch that can cause errors to compound over a generated sequence. We investigate generative adversarial fine-tuning as a means of introducing autoregressively generated response segments into language-model training, exposing the model to its own predictions during training. To permit gradient propagation to the generator, we leverage Gumbel-Softmax sampling rather than standard discrete token sampling to obtain a differentiable next-token representation, allowing generative adversarial network-based training. Using a convolutional discriminator, which distinguishes generated responses from reference ones, and a transformer language model as the generator, we attempt to improve training stability and response quality through the use of regularization techniques. In particular, during training we progressively unfreeze the discriminator's convolutional layers and optimize the generator using a weighted combination of adversarial and standard supervised fine-tuning objectives. Using an LLM-based evaluator, our best anchored configuration improves over the supervised fine-tuning baseline in all three evaluated categories, including relative gains of 9.9$ percent in Summarization and Science, respectively. These results provide preliminary evidence that the adversarial generation objective can complement supervised fine-tuning in the studied setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.