acceptodds
Under review as a conference paper at ICLR 2027

Understanding Context Quality and Improving Batch Retrieval for Tabular Foundation Models

Abstract

Tabular foundation models predict new samples from labeled reference examples without task-specific model training. Using a separate context for each query or query cluster incurs substantial computational cost, motivating the selection of an effective shared context for an entire query batch under a limited budget. Existing retrieval methods typically use fixed selection rules whose performance varies substantially across tasks. Through exhaustive synthetic experiments and controlled analyses on real datasets, we find that a sample's predictive contribution is closely related to local class support, support accumulation, allocation across queries, and class composition. Based on these observations, we propose Preference-Aware Integrated Retrieval (PAIR), an adjustable retrieval framework that combines support accumulation, coverage allocation across queries, and class reservation, with configurations selected through validation. Its coverage objective evaluates candidates according to existing coverage and exhibits diminishing returns, enabling efficient greedy optimization with an approximation guarantee. Across 51 classification tasks, PAIR achieves the best mean rank of 1.91 among seven methods, with its untuned default ranking second, and the gains persist across four frozen models, multiple context budgets, and regression tasks. Its mean inference time is approximately one twenty-fifth that of per-query retrieval, with corresponding differences in predictive quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.