acceptodds
Under review as a conference paper at ICLR 2027

One Table, Many Queries: Benchmarking Generalist Tabular Models across Generation, Imputation, and Prediction

Abstract

Tabular generation, missing-value imputation, and supervised prediction are usually developed and benchmarked as separate problems, even though each can be written as a conditional query to the same table distribution. This fragmentation leaves a basic question unresolved: does competence on one query carry to others, or does a multi-interface model merely expose several APIs? We introduce OneTableBench, to our knowledge the first benchmark designed to turn this shared probabilistic view into a unified empirical study. Rather than merging three leaderboards, OneTableBench preserves task-specific data, information, and evaluation contracts while aligning 56 datasets, 46 registered methods, and 20 semantic capability axes. Its evaluation forms a four-level evidence chain: auditing metric redundancy, reliability, and construct validity; measuring cross-task rank alignment on matched support; estimating transfer with matched interventions over all seven non-empty combinations of generation, imputation, and prediction objectives; and comparing fixed model identities with deployable specialist portfolios through Pareto and leave-one-dataset-out analyses. The resulting picture is neither universal transfer nor complete task independence. Metrics are compressible but remain multidimensional, and method rankings agree only along particular capability axes. Observational alignment does not produce a robust off-diagonal transfer effect under our prespecified gate; several joint-objective interactions are instead sub-additive, and classification and regression induce different objective frontiers. Deployment conclusions are similarly profile-dependent: evaluated generalists can surpass selected specialist portfolios on downstream utility, whereas the portfolios retain a clear advantage on categorical fidelity. OneTableBench therefore reframes tabular generality from binary task support or a single leaderboard rank into a capability profile over queries, contracts, and deployment priorities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.