acceptodds
Under review as a conference paper at ICLR 2027

Benchmaxxing: Selecting Data that Targets Many Benchmarks with Scalable Gradient-Based Methods

Abstract

Data quality is multi-dimensional: whether a sample is “good” depends on the model's downstream use, the checkpoint's weaknesses, and the composition of the surrounding dataset. Mainstay methods rely on human-like heuristics that, while useful, cannot target specialized use cases in a principled way. We advocate for benchmark-targeted data filtering, which scores samples based on their predicted effect on downstream performance. Prior approaches have been computationally expensive and limited to older models and small training scales. Our method, Benchmax, utilizes coverage selection criteria that improve the diversity and effectiveness of selected data, and combines gradient sketching with fast KNN search for efficient selection at scale. We distill the selection criteria into a multi-head classifier that simultaneously predicts a sample's impact across many benchmarks. This low-cost classifier scales to modern training regimes, scoring 1 trillion tokens for approximately $500 in cloud costs, making principled data selection practical.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.