acceptodds
Under review as a conference paper at ICLR 2027

Beyond Group Proportions: Quality Filtering and Dialect Gaps

Abstract

Quality filters decide which documents survive into a language model's training data, and their effects fall unevenly across varieties of English. The usual way to check for this is to count how much data each group keeps. We ask whether that count tells the right story for African American English (AAE), the variety spoken by many Black Americans, by comparing models trained on datasets that contain exactly the same number of AAE-aligned and white-aligned documents but different documents within each group. In a controlled continued-pretraining experiment on tweets, replacing the documents chosen by the DataComp-LM quality classifier with different documents at the same group counts removes 63% of the widening in the likelihood gap between the groups relative to uniform selection, and it improves fit to both groups relative to the filter's own documents. Restoring the original group share also narrows the gap, so the count matters, but it is not the whole account. In web pretraining from scratch at a matched token budget, an additional cut with the same classifier on an already-curated corpus raises the between-group perplexity ratio by 2.9%, whereas a second pool and budget give an estimate whose interval includes zero. Reweighting documents kept by a perplexity filter does not reproduce the benefit of changing them. These results argue for auditing which examples of a variety remain in the training data, and not only how many.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.