Measuring Structure in Tabular Foundation Model Priors
Abstract
Tabular Foundation Models (TFMs) are trained to predict labels given in-context datasets and queries sampled from tasks (data-generating processes). These tasks are sampled from diverse but synthetic priors. The prior determines how often different tasks appear during pretraining, hence the finite computation spent training on them. However, it is unclear from the various choices in designing a prior, how much compute to allocate to each task. Motivated by the notion of epiplexity, we characterise the learnable predictive structure accessible to a model to help better allocate compute across tasks. To do this, we introduce task expected predictive information gain (TEPIG), a model-dependent measure of the expected reduction in joint predictive log loss obtained by observing training datapoints from a task. We illustrate its behaviour in controlled examples and test the hypothesis that TEPIG, when measured with cheaper proxy models, can guide the allocation of pretraining compute for TFMs. In a tractable Gaussian process setting, TEPIG-guided sampling improves predictive accuracy and convergence when evaluated under the original prior. On more complex priors, including a state-of-the-art prior used for open tabular foundation models, it allows for training in fewer steps. These findings demonstrate TEPIG’s practical value for measuring model-dependent structure in tabular foundation model priors and suggest its potential as a tool for more efficient compute allocation for TFM training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.