acceptodds
Under review as a conference paper at ICLR 2027

StruggleTab: A Challenging Benchmark for Tabular Synthesis Models

Abstract

While recent tabular synthesis models have achieved remarkable performance, the standard framework for evaluating these models has remained too simple. In fact, the use of overly-sanitized evaluation datasets has created a misleading appearance that synthetic tabular data is nearly at parity with real-world data. This misguided impression quickly falls apart when models are evaluated on more challenging datasets. We introduce the *StruggleTab benchmark*, comprising seven datasets containing three challenging characteristics common in real-world data: (1) text, (2) feature imbalance and (3) strict invariants. Our audit of 5 leading tabular synthesis models (TabDLM, Tabby, TabDiff, ICL, CTGANP) on *StruggleTab evaluation metrics* reveals many failure modes including high dimensionality, sequence lengths and more. Concerning results include easy distinguishability from non-synthetics (average % score on C2ST) and all 5 models receiving a negative score in the standard downstream utility metric (Machine Learning Efficacy) on the financial Limit Order Book task. StruggleTab raises the empirical standards for tabular synthesis, highlighting the need for architectures compatible with real-world heterogeneity.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.