acceptodds
Under review as a conference paper at ICLR 2027

TabTreeVFM: Endpoint Variational Flow Matching with Boosted Trees for Mixed-Type Tabular Generation

Abstract

Synthetic tabular data supports downstream analytics, machine learning, and data sharing when real data is scarce, sensitive, or costly to collect. Mixed-type synthesis must capture joint dependencies while respecting heterogeneous column structure, yet high-fidelity generators can reproduce training records rather than learn the underlying distribution. Gradient-boosted trees suit such heterogeneous data, but existing tree-based flow models regress ambient-space velocities, imposing a Euclidean geometry on inherently multinomial categorical variables. To address this, we propose TabTreeVFM, a tree-boosted endpoint-posterior model with Gaussian regression heads for continuous columns and multinomial classifier heads for categorical variables, in place of direct velocity regression. The resulting posterior means recover the flow-matching velocity analytically, and training decomposes into independent XGBoost subproblems across timesteps and feature groups. Across 10 standard mixed-type benchmarks, TabTreeVFM attains the best overall average rank among 12 generators across fidelity, utility, and empirical privacy. Moreover, we introduce Mem, a held-out-calibrated measure of excess full-row reproduction, and find that several strong neural baselines reproduce training rows substantially above the natural collision rate, while TabTreeVFM remains close to the held-out reference. Together, these results show that type-aware endpoint prediction combines the tabular inductive bias of boosted trees with principled flow-matching geometry, yielding strong fidelity and utility alongside a favorable empirical privacy and memorization profile.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.