acceptodds
Under review as a conference paper at ICLR 2027

Back to the roots: revisiting tree-based autoregressive models for tabular synthesis

Abstract

We challenge the hegemony of deep generative models, especially diffusion models, for mixed-type tabular data generation. We revisit a popular method for mixed-type tabular synthesis from the statistical community – tree-based autoregressive models – and target two existing issues: (1) dependence on a single autoregressive factorization, and (2) using the same data to learn tree partitions and perform sampling at inference-time. Our method (Synthpop+) forms a mixture over random-order tree-based autoregressive models and uses out-of-bag observations for honest leaf-value estimation. Across 16 datasets and 11 modern baselines, Synthpop+ scores best in terms of multivariate fidelity: it decreases the error of a classifier 2-sample test by 49% w.r.t. the best-performing baseline (TabCascade, a recent cascaded diffusion model for tabular data), while simultaneously decreasing overfitting risk (DCR Share) and membership disclosure risk. We achieve this performance gain at 4% of the training time while training on consumer-grade CPUs instead of GPUs. These results indicate that these relatively simpler classical methods still comfortably beat deep generative models when given enough attention; we challenge practitioners to reconsider when deep generative models are actually required for tabular synthesis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.