acceptodds
Under review as a conference paper at ICLR 2027

TabBench-Bio: A Living Benchmark for Machine Learning on High-Dimensional Biomedical Tables

Abstract

Biomedical tables often combine thousands of measured variables with only tens or hundreds of labelled samples, a regime that is poorly represented in general-purpose tabular benchmarks. We introduce TabBench-Bio, a living and interactive benchmark of 43 biomedical datasets spanning multiple domains. Under a shared cross-validation protocol, we compare classical estimators, neural networks, and tabular foundation models across 28 feature-by-sample operating points. At the reference cell of 10,000 features and 100 training samples, RealTabPFN 2.5 has the highest point estimate, closely followed by TabPFN 3 and Logistic Regression. A paired bootstrap over the target pool separates RealTabPFN 2.5 from TabPFN 3 by 87 Elo (95% interval [43, 129]). Tabular foundation models generally occupy the leading ranks, while the strongest configuration depends on the operating point and biomedical modality. The AutoML framework AutoGluon, using its one-hour “extreme” preset, is configured as a separate resource-intensive reference and reported here at the reference and full cell. Fold-level predictions, run status, and deterministic aggregations make every reported result reproducible and reusable. The benchmark is open to contributions of new biomedical datasets. The interactive leaderboard is available at https://anonymous_iclr.codeberg.page/TabBench-Bio-Anonymous/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.