Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage
Abstract
In verifiable domains, such as math and coding, where multiple solutions can be generated in parallel and scored by a verifier, getting just one correct solution among many attempts can be more important than the individual pass rate of each attempt, especially for difficult tasks where a correct solution might be a needle-in-the-haystack. Post-training methods such as supervised fine-tuning (SFT) and reinforcement learning based preference fine-tuning in Large Language Models (LLMs) have been shown to concentrate their outputs around a few modes. A common method to increase LLM output diversity is to increase a “temperature” parameter, but the effectiveness of this has been shown to be limited (Banayeeanzade et al., 2026; Holtzman et al., 2020). To address this issue we introduce Don't Repeat Yourself Supervised Fine-Tuning (DRY-SFT): a post-training method that increases output diversity and improves coverage — coverage being how likely it is for there to be at least one correct solution among many generated attempts. DRY-SFT has two simple stages: first, for each problem, sequentially generate K solutions, where, for each subsequent generation, the prompt contains the problem plus all of the prior problem attempts and asks for a different solution. Second, fine-tune the LLM on each solution attempt as if it was generated independently by stripping away the prior attempts from the context. This process involves no reward, verifier, or correctness filter. The result is a model that is trained to produce many diverse solutions. We evaluate DRY-SFT on three coding benchmarks and show that it raises pass@100+ by 10.8, 12.5, and 12.4 points across HumanEval+, MBPP+, and DS-1000 at a small cost to pass@1+. We also measure the structural diversity using abstract syntax tree (AST) edit distance of passing solutions and find that this rises significantly across all three benchmarks. Additionally, DRY-SFT solves 244 of 600 problems the base model was not capable of solving in the same 200 attempts. Finally, across nine open-weight models, lower structural diversity of the base model significantly predicts higher DRY-SFT performance gains, indicating that our method is especially effective on more mode-collapsed LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.