Is More Always Better? Unveiling the Impact of Pre-training Iterations on Fine-tuning Performance in Text-to-Image Models
Abstract
As pre-training progresses, checkpoints at different stages may adapt to downstream tasks at different rates. Selecting the optimal pre-trained checkpoint is crucial for efficient downstream adaptation. Conventional wisdom suggests that more pre-training generally leads to better downstream performance, while findings from cognitive science suggest that younger learners often adapt more quickly. Motivated by this contrast, we investigate whether text-to-image (T2I) models exhibit a similar "youth advantage" during fine-tuning. Understanding this phenomenon could enable more compute-efficient adaptation strategies. Through experiments on two T2I backbones and three downstream datasets across ten pre-trained checkpoints, we reveal that when fine-tuning tasks conflict with the pre-training data, moderately pre-trained ("Youth") models adapt most quickly in full fine-tuning, whereas under LoRA, the most extensively pre-trained ("Expert") models converge the fastest. We explain this discrepancy with a fluid/crystallized intelligence framework from cognitive science, suggesting that full fine-tuning benefits from the fluid intelligence of young models, whereas LoRA leverages the crystallized intelligence of mature models. Our findings offer a unique perspective on practical checkpoint selection, enabling computational savings compared to using the most pre-trained model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.