DETECT LLM FINE-TUNING OVERFITTING EARLY
Abstract
Four epochs out of fifty decide whether a fine-tuned model overfits, and four epochs are enough to catch it. Across twenty-two matched-pair comparisons spanning 3.8B to 32B parameters, an early check at epoch 4 calls the final test-set gap for every readable run, with the remaining pair settling by epoch 10. Both thresholds were locked in advance. The results uncover a clear pattern: models distilled from reasoning checkpoints overfit downstream tasks more than their standard instruction-tuned twins. Across four base models they show 1.06 to 4.60x larger drift and 1.05 to 1.84x wider final test gaps. In contrast, models built directly by vendors for reasoning or domain tasks behave much like their standard twins. Pairing models from the same base keeps the comparison clean by holding architecture, vocabulary and size constant. Using only raw training and validation loss curves, our metric splits the gap into two plain numbers: roughness a, how far the training losses sit from a smooth descent, and drift c, how far the validation curve pulls away. Their ratio c/a pinpoints what kind of overfitting is happening. On an unseen 32B pair run across two seeds, the rule flags heavy overfitting at epoch 4, correctly anticipating final drifts over 4.4x. While standard loss gaps top out at sixteen correct calls even with best-case hindsight tuning, our early rule scores a perfect eighteen for eighteen.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.