acceptodds
Under review as a conference paper at ICLR 2027

Dataset Alignment Over Diversity

Abstract

Scaling pre-training data volume has become the industry standard for improving language model performance, yet this approach is computationally wasteful when the goal is strong downstream performance on a specific task, especially when models are often pre-trained on diverse datasets rather than specialized training on datasets with overlapping domain coverage. We challenge this paradigm and propose Dataset Alignment Over Diversity (DAOD), a less compute dependent framework that predicts how well a pre-training dataset will transfer to a target downstream task without requiring any model training. DAOD scores candidates using a closed-form function of three geometric properties extracted from a randomly initialized model: centroid alignment between candidate and target embeddings, alignment friction (variance), and intra-dataset diversity. Across a benchmark of 8 candidate datasets evaluated against a mathematical reasoning target (OpenWebMath) and a second target validated across three architectures (Gemma-3 1B, LLaMA-3.2 1B, LFM-1.2B), DAOD reliably sorts candidates into the correct pre-training tier, capturing 95–98% of the achievable downstream performance gain while using only 25–38% of the total pre-training compute. The underlying score also shows strong rank correlation with ground-truth fine-tuning loss (maximum Spearman = 0.93 on OpenWebMath), though we note this correlation is computed over a modest number of candidates and is best interpreted alongside the tiering results, which are more robust to individual rank swaps. Together these results establish that data alignment, not volume, is the primary driver of efficient transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.