Benchmaxxing: Selecting Data that Targets Many Benchmarks with Scalable Gradient-Based Methods
Abstract
Data quality is multi-dimensional: whether a sample is “good” depends on the model's downstream use, the checkpoint's weaknesses, and the composition of the surrounding dataset. Mainstay methods rely on human-like heuristics that, while useful, cannot target specialized use cases in a principled way. We advocate for benchmark-targeted data filtering, which scores samples based on their predicted effect on downstream performance. Prior approaches have been computationally expensive and limited to older models and small training scales. Our method, Benchmax, utilizes coverage selection criteria that improve the diversity and effectiveness of selected data, and combines gradient sketching with fast KNN search for efficient selection at scale. We distill the selection criteria into a multi-head classifier that simultaneously predicts a sample's impact across many benchmarks. This low-cost classifier scales to modern training regimes, scoring 1 trillion tokens for approximately $500 in cloud costs, making principled data selection practical.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.