AutoTabBench: Automated Tabular Benchmark for Tabular Foundation Models with Evidence-Grounded Scope Analysis
Abstract
Tabular foundation models (TFMs), pretrained on large-scale tabular data, have emerged as powerful data-analysis assistants for real-world applications, supporting data interpretation, prediction, and decision-making across diverse domains. As both general-purpose and domain-specific TFMs continue to grow, a natural question arises: how can we comprehensively and systematically evaluate their capabilities and determine which models are best suited to different tasks and conditions? Existing benchmarks provide the main basis for such comparisons, but they often reduce model evaluation to aggregate predictive performance over limited conditions. As a result, they provide limited evidence about how model advantages change across task types, distribution shifts, data characteristics, and perturbations. We argue that reliable evaluation requires making explicit the boundary of conditions under which a benchmark supports its conclusions. To this end, we introduce BenchScope, an evidence-grounded framework for benchmark scope analysis and design. BenchScope recovers verifiable evaluation evidence from papers, code, configurations, and manifests, and structures it across task, generalization, and protocol dimensions. Using explicit evidence requirements and decision rules, it characterizes benchmark boundaries by identifying covered conditions, missing coverage, protocol constraints, and unresolved evidence, and translates these findings into benchmark-construction requirements. Across six benchmarks, structured boundary recovery improves from 80.56% to 94.44%. Applying BenchScope to 37 benchmark entries, we construct AutoTabBench, an 80-task automated tabular benchmark covering routine prediction, natural distribution shifts, and controlled stress conditions, under which model advantages vary substantially across evaluation settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.