acceptodds
Under review as a conference paper at ICLR 2027

CleanCore: Data-Quality-Aware Coreset Selection for Learning over Dirty Tabular Data

Abstract

Multiple error types and high training costs in large-scale tabular data make it difficult to balance predictive performance and runtime efficiency. Existing data cleaning, coreset selection, and hybrid methods typically focus on one objective or specific error types. To address this, we propose **CleanCore**, a data-quality-aware coreset selection method that handles multiple types of errors. The method iteratively uses training signals for error typing and lightweight handling, and dynamically maintains training coresets based on sample reliability. Theoretical analysis provides support from three perspectives: error-typing reliability, error-handling bias, and coreset representativeness. Comparisons with eleven baselines on eleven public tabular datasets using MLP and tabular ResNet show that **CleanCore** achieves the best ACC and F1 in 21 and 20 of the 22 model–dataset settings, respectively, delivers approximately – runtime speedups over three data cleaning methods, and maintains strong performance under real-world dirty-data conditions and different injected-error settings. Further experiments validate its stability and the effectiveness and efficiency of its main mechanisms.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.