Ambient LLMs: Data-Efficient Supervised Finetuning via Temperature Matching
Abstract
Collecting large-scale, diverse, and high-quality post-training datasets for autoregressive language models is challenging. Consequently, Supervised Fine-Tuning (SFT) is typically applied to datasets significantly smaller than those used for pre-training, which leads to overfitting and the catastrophic forgetting of the tails of the data distribution. To prevent this collapse, we introduce \methodname, a principled framework for leveraging imperfect predictions to augment the fine-tuning set. In our method, a “frozen” version of the pre-trained model serves as a noisy teacher that the fine-tuned model must match at a coarse level. In particular, we require that the high-temperature versions of the conditional distributions produced by the post-training model should match the high-temperature versions of the pre-training ones. The temperature hyperparameter controls the level of trust in the pre-trained model, with the limit reverting back to regular SFT and to knowledge distillation. We demonstrate that our approach outperforms standard SFT across domain adaptation, instruction-tuning, and reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.