acceptodds
Under review as a conference paper at ICLR 2027

ATCS: Accelerating Target-model-aware Coreset Selection

Abstract

Large-scale supervised fine-tuning (SFT) of foundation models increasingly relies on instruction data pools aggregated from heterogeneous sources such as the web, open communities, and synthetic generation pipelines. To support reliable model adaptation, such data must be transformed into a compact, high-value training set by removing redundant, noisy, and low-utility examples. This transformation naturally calls for data selection methods that can reduce the training pool while preserving the information needed for downstream adaptations. Target-model-aware coreset selection is particularly effective for this purpose because it evaluates candidates according to their utility for the model to be fine-tuned. However, existing target-model-aware methods are difficult to deploy at scale. Specifically, they typically require the target model to score every candidate instruction-response pair, turning data selection itself into a costly inference workload. To address this bottleneck, we propose ATCS, a general acceleration framework that treats selection as a two-stage data processing pipeline. ATCS exploits the agreement between the proxy and target models in identifying low-value samples and salient instructional content, reducing both the number and length of records evaluated by the target model. Besides, ATCS is compatible with multiple utility functions and supports cross-utility selection across coarse filtering and target-model reranking. Experiments on large-scale instruction-tuning datasets show that ATCS preserves downstream performance comparable to exhaustive target-model-aware selection while reducing end-to-end selection time by more than 77%, making model-aware coreset selection practical for scalable LLM data preparation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.