TabClean: Evidence-Guided Program Synthesis for Scalable Tabular Data Cleaning
Abstract
Reliable analytics and machine-learning pipelines depend on clean tabular data, yet production tables contain heterogeneous errors that require different repair mechanisms. Constraint-based systems require explicit rules, while learning-based systems require labels or training. LLMs can help infer repairs, but repeatedly reasoning over cells or chunks incurs inference costs as tables grow and new batches arrive. Synthesizing executable cleaning logic offers a way to amortize this reasoning, provided that the evidence supports both the transformation and the conditions under which it should apply. We present TabClean, a model-training-free system for scalable tabular data cleaning through evidence-guided program synthesis. TabClean uses table evidence and execution feedback to synthesize guarded cleaning programs, amortizing synthesis-time LLM cost through deterministic execution over large tables and compatible batches. Table profiles and annotated examples guide the choice of repair mechanisms, target representations, and guard predicates. Cell-level feedback from executed candidates then guides program revision, and a deterministic controller retains the best development-set program. Across six benchmarks, our approach achieves a mean F1 of 0.842, improving on two LLM baselines by 0.456–0.764 with identical development examples and test rows. On a 200,000-row table, it achieves 0.983 F1 in 4.42 minutes, an approximately 1,200× speedup over a semi-supervised repair baseline. Its mean LLM API cost across the six benchmarks is 98% lower than that of an iterative LLM repair baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.