acceptodds
Under review as a conference paper at ICLR 2027

TabQueryBench: A Query-Centric Benchmark for Synthetic Tabular Data

Abstract

Synthetic tabular data support use cases like data sharing, model development under access restrictions, and rapid prototyping of analytical workflows. Modern generative models are evaluated by their statistical similarity, correlation structure, privacy, and downstream machine-learning utility. However, such evaluations leave a gap: they rarely test the structure that matters for analytical queries. We present TabQueryBench, a query-centric benchmark for evaluating synthetic tabular data with SQL-shaped analytical queries. TabQueryBench distills 12 public query sources into 44 reusable query templates and grounds them to each dataset through a policy-guided template-to-SQL pipeline. Across 49 datasets and 11 generative models, it produces more than 100 executable queries per dataset while preserving cross-model comparability. Our evaluation reveals a substantial gap between conventional and query-centric fidelity. Even the strongest model, RealTabFormer, reaches only query-centric fidelity compared with for real data. Failures are most pronounced for high-cardinality attributes and local or tail queries, while model quality also trades off against generation cost. These results show that synthetic data can appear statistically realistic yet still fail to preserve the analytical behavior of real data, motivating query-centric fidelity as a necessary complement to existing evaluation metrics. TabQueryBench is open-sourced and available for anonymous review.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.