acceptodds
Under review as a conference paper at ICLR 2027

FAITH: Typed Cell-Level Tabular Generation with a Frozen Language Backbone

Abstract

Synthetic tabular data must reproduce heterogeneous column domains, their dependencies, and the predictive structure carried by complete rows. Pretrained causal language models offer a reusable mechanism for composing contextual dependencies, but existing language-model generators usually retain a textual interface: a cell expands into several tokens, prediction is performed over a general vocabulary, and numerical values are emitted as strings. We introduce FAITH, which keeps a frozen causal backbone while replacing its language interface with typed cell encoders and schema-local output distributions. Every cell occupies one causal position; categorical columns are predicted on their observed support, and numerical columns combine quantile-bin selection with a continuous conditional Beta density. This construction separates the sequence model that shares context across columns from the distributions that define individual columns. Across ten datasets and five baselines, FAITH obtains the highest primary synthetic-train-to-real-test utility on four datasets and its strongest joint quality on the dependency-intensive Allstate and HBA1C cases. On these datasets, FAITH trains - and samples - faster than the token-level GReaT baseline. The evaluation separately reports fidelity, utility, runtime, and empirical disclosure behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.