Quality Is Not One Thing: Auditing, Distilling, and Blending LLM-Judge Annotations for Pre-Training Data Selection
Abstract
Curation of data in language model training commits the corpus to particular notions of 'quality'. We study that commitment in three parts: judging different notions of quality, optimising them in a mix on a probability simplex, and training and evaluating on the filtered corpora. An LLM-judge scores a fixed one-million-document English Common Crawl sample using 22 unique prompts, and we distil each set of scores into lightweight classifiers. We train three classifier types (fastText, a frozen embedding with an MLP head, and fine-tuned BERT) for a total of 192 models. The scores share a dominant content factor (first principal component of variance) but can also be very distinct, e.g., conversationality is largely uncorrelated with the other tested axes. Disagreement between scores is largest among the documents different classifiers would label `high quality'. Within-document spread across standardised scores rises from in the bottom decile of mean score to in the top, and classifiers with similar aggregate metrics keep different top- sets (Jaccard –). We optimise a weighted mix of five of these scores on a probability simplex. We show a weighted sum of the raw scores performs worse than a corpus selected only with the Dolma 3 quality filter. However, putting each score on a shared rank scale and keeping documents above a Dolma 3 classifier cutoff lets those dimensions combine more effectively. Training on this latter mix, the Dolma 3 classifier filtered data, and a random crawl sample in a B dense model and a B-AB mixture-of-experts model shows the the raw weighted sum does not beat Dolma 3 classified, while the rank-scaled mix does and also has some positive transference to performance outside of tested English language capabilities (including coding, mathematics, and German language).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.