Beyond Oracle Headroom: Evaluating Context Selection in Tabular Foundation Models
Abstract
Better contexts in hindsight need not yield better decisions at deployment. We study this gap in tabular foundation models using TabPFN and TabICL. Development evaluations show mean Raw Oracle Headroom of 0.318 in negative log-likelihood (NLL), yet expanding a nested candidate bank increases headroom without corresponding gains for the tested selectors. An exact decomposition separates fixed candidate-rule advantage from observed per-query switching opportunity and shows that similar headroom can arise from different candidate-loss structures across datasets. On ten prospectively selected holdout datasets, simple geometric pre-outcome policies beat uniform random choice by 5.69–8.89 percentage points among bank-positive queries. However, post-hoc frequency matching and a direct P2 fixed-rule comparison leave added rule-identity assignment value unestablished. P2 also has exact score ties on 99.80% of queries: tied contexts often share score-determining bottlenecks yet yield different predictions, and its lexicographic tie-break outperforms uniform choice among minima on both models. Once bank predictions are available, predictive centrality (PC) improves over CRUMB in the original formal test, while an advantage over same-bank averaging is not established. On the fresh holdout, gains over averaging are supported only for TabICL, and neither backbone establishes an advantage over CRUMB. Mean gains can also coexist with losses on many individual queries. Together, these results separate hindsight opportunity, policy-level success, and deployment value, showing that context-selection claims depend critically on the comparator and decision stage.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.