Advancement or Acceleration? How Data Curation Improves LLM Training
Abstract
While data curation is widely credited with improving large language model (LLM) training, a basic question remains unresolved: does curation expand what a model can ultimately learn, or merely reach the same solution faster? In this paper, we formalize this distinction between *advancement* (improving the attainable performance) and *acceleration* (reaching the same performance with less data). Through a closed-form analysis and real-world LLM experiments, we show that curation can advance a model only when a train–validation gap is present. Without such a gap, it can only accelerate, and no allocation beats the uncurated baseline. This reframes advancement as a property of the curation *task* rather than the curation method. Under a common gap-centric view of data mixture and selection, we operationalize this condition as the -test (), an observable indicator computable prior to expensive curation. Across extensive mixture and selection experiments, the -test successfully predicts when curation helps and reveals that many reported gains are merely acceleration. Ultimately, our findings show that the true ceiling of data curation is dictated by validation construction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.