acceptodds
Under review as a conference paper at ICLR 2027

Self-Generative Adversarial Fine-Tuning for Large Language Models

Abstract

Post-training is central to making large language models (LLMs) useful, but it relies on high-quality examples and feedback that are scarce and costly to collect. Iterative self-improvement makes better use of the available data by generating new training examples, provided the model can assess what it generates. Existing methods accept model-generated data without a grounded correction signal or rely on self-evaluation scores that may bias toward the model's own style. We introduce Self-Generative Adversarial fine-tuning for LLM (SGALM), which learns that assessment from real examples. At each round, one LLM generates complete examples from few-shot prompts and learns to distinguish its generations from the real data. The generation is rewarded to be realistic while the discrimination is trained by the truth, instantiating a generative adversarial game without a external feedback or new human labels. Across tasks from different domains, SGALM significantly improves over SFT and existing self-improvement method on different backbones. SGALM also serves as a synthetic-data engine and a realness-based filter, whose generated data continues to improve downstream training. We also diagnose discriminator score and generation diversity which provide evidence against self-evaluation bias and mode collapse.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.