Standard Data Scaling Analyses Obscure Loss of Important Information in Math Reasoning
Abstract
In deep learning, increasing dataset size has been shown to improve the performance of deep neural networks. However, it is unclear if adding more data negatively impacts certain abilities learned from existing data, obscured by an overall net increase in performance. Understanding this is especially important in the current large language model era, where data scarcity has become a pressing issue. We discover that when performing fine-tuning on mathematical reasoning tasks, adding more training data causes the model to incorrectly answer a large portion of previously correctly answered test samples. This remains true even with popular test-time scaling techniques, which can iron out inconsistencies in model predictions. To better understand this phenomenon, we show that large language models fine-tuned on the same data learn very different functions across different random seeds, exhibiting extremely high predictive multiplicity. This work contains novel insights that encourage practitioners to investigate if additional data is yielding equitable benefits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.