Dimension-Free Data Filtering
Abstract
Selecting high quality data is a crucial prerequisite for training and fine-tuning large language models. While empirical filtering heuristics are widely used, they often lack rigorous guarantees regarding how well the filtered data aligns with the target task. In this paper we propose a theoretical framework for data filtering that provides explicit distribution matching guarantees. Given a small set of high quality examples and a larger set of lower quality examples, our goal is to design a filter to identify more high quality examples. Our main result is a simple mechanism that generates a distribution of examples that is within total variation distance of the high quality distribution. We also show nearly matching information-theoretic lower bounds.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.