Fine-Tuning Fixes Mode Collapse and Over-Dispersion in LLMs
Abstract
Recent work raises concerns that outputs from large language models (LLMs) tend to be under-diverse: they repeat or resemble one another more often than responses from the population they are meant to represent, a phenomenon known as mode collapse. In this work, we show that whether mode-collapse—or its opposite—occurs depends on the specific model and dataset used. Further, with sufficient supervised fine-tuning (SFT) data, LLM output diversity converges toward that of the target distribution from which fine-tuning data are sampled. To quantify this comparison, we measure the probability that two responses sampled independently conditional on the same fixed prompt coincide (collide). In one set of experiments, we use exact token sequences; in the other, we use the generalised version of that—the expected similarity between responses under a kernel. We define miscalibration as a nonzero model-minus-target collision gap. We derive a bias–variance decomposition of the expected gap between the model's and target's collision probabilities, showing that SFT is not inherently biased toward mode collapse or its opposite: finite-sample SFT can leave a model either under- or over-dispersed, depending on the model and dataset. Finally, we show that the absolute gap is bounded by the square root of the Kullback–Leibler (KL) divergence from the target distribution to the model. Consequently, a model sufficiently close to optimal under population cross-entropy cannot exhibit arbitrarily miscalibrated diversity. We test the decomposition and the bound in three experiments: (1) we fit 100 small transformers at each of 100 log-spaced sample sizes on each of two synthetic order-16 languages; (2) we fine-tune four LLMs using low-rank adaptation on responses from the General Social Survey, American National Election Studies, and World Values Survey; and (3) we repeat Experiment (2) on CodeNet, a dataset of human solutions to coding tasks, measuring program similarity with normalized Zhang–Shasha edit distance between canonical abstract syntax trees. We find substantial heterogeneity in over- and under-diversity across models and datasets. More target data moves model diversity toward the human (or synthetic target) level in all experiments, consistent with our theoretical predictions. These results show that diversity miscalibration can arise from finite-sample error and shrink as SFT better approximates the target distribution. Accordingly, as the sample size of target-distribution data increases, model diversity moves toward the target level.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.