ACE: Lifting Data Attribution from Samples to Clusters via Intrinsic Gradient Similarity
Abstract
Data quality has become increasingly central to language model training, making effective data selection important. Yet many influential data-selection methods score each sample independently at high cost. Such pointwise scoring overlooks similarities in gradient behavior across samples, leading to limited scalability and redundant selection. To address these limitations, we introduce **ACE** (Data **A**ttribution via **C**luster-Level **E**stimation), a scalable framework that exploits intrinsic gradient similarity to lift both utility estimation and budget allocation from individual samples to clusters. ACE groups samples into clusters by gradient similarity, scores a few representatives per cluster to estimate cluster-level utility, and allocates the selection budget across clusters. We prove that the resulting estimator is unbiased and has controlled variance, providing a principled basis for reliable cluster-level attribution. Across post-training experiments, ACE achieves an end-to-end speedup of up to **33.8×** and consistently outperforms its sample-level counterparts. Using only **10%** of the data, it even surpasses **full-data** training on Llama-3.2-3B and Mistral-7B. Pre-training experiments further validate its scalability to larger data pools: ACE improves upon sample-level OPUS, achieves an **8.86×** speedup, and delivers the strongest overall results among the baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.