Token Mass, Not Geometry: A Confound in Embedding-Based Data Selection for Language Model Pretraining
Abstract
Cluster/embedding-based data-selection methods choose pretraining subsets by geometry: representativeness (proximity to a cluster centroid), atypicality, or redundancy. We stress-test this family under matched training compute, three seeds, a common held-out evaluation, and an explicit random baseline. Under the usual *fixed-document-budget* protocol, near-centroid (“representative”) selection beats random on out-of-distribution (OOD) loss by a small, seed-stable margin (0.01 nats on C4). We show this gain is a **retained-token-mass confound**, not a property of embedding geometry. Centroid distance is strongly anti-correlated with document length ( in an LSA space, in the OPT-125M space), so at a fixed document budget, representative selection keeps longer documents and hence *more tokens*, training with less repetition. When we instead match the *token* budget across arms, the benefit vanishes: near-centroid selection ties or slightly underperforms random. This holds across model scales (124M and 350M, three seeds each) and two OOD sets, and a progression of controls shows the near/far-centroid gap shrinking from nats (document budget) to (length-stratified) to (token-matched). Native SemDeDup, evaluated in its own OPT-125M embedding space, also does not beat random on this corpus. We recommend that evaluations of embedding-based selection control for retained token mass and document length, and report against a random baseline at matched compute with multiple seeds: practices that, in our experiments, overturn an apparently positive result.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.