Beyond Group Proportions: Quality Filtering and Dialect Gaps
Abstract
Quality filters decide which documents survive into a language model's training data, and their effects fall unevenly across varieties of English. The usual way to check for this is to count how much data each group keeps. We ask whether that count tells the right story for African American English (AAE), the variety spoken by many Black Americans, by comparing models trained on datasets that contain exactly the same number of AAE-aligned and white-aligned documents but different documents within each group. In a controlled continued-pretraining experiment on tweets, replacing the documents chosen by the DataComp-LM quality classifier with different documents at the same group counts removes 63% of the widening in the likelihood gap between the groups relative to uniform selection, and it improves fit to both groups relative to the filter's own documents. Restoring the original group share also narrows the gap, so the count matters, but it is not the whole account. In web pretraining from scratch at a matched token budget, an additional cut with the same classifier on an already-curated corpus raises the between-group perplexity ratio by 2.9%, whereas a second pool and budget give an estimate whose interval includes zero. Reweighting documents kept by a perplexity filter does not reproduce the benefit of changing them. These results argue for auditing which examples of a variety remain in the training data, and not only how many.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.