acceptodds
Under review as a conference paper at ICLR 2027

Tabular Selective Attention

Abstract

Tabular foundation models (TFMs) predict each test row by attending over a context of labelled training rows, so a prediction is only as good as the set of rows attention averages over. While these models perform well in low-data regimes, current state-of-the-art models employ context curation to generalize over larger datasets. Existing curation methods either sharpen attention with a temperature or shrink the context with nearest-neighbour retrieval. We prove that temperature scaling faces an inherent tradeoff: a single scalar must balance suppressing the uninformative tail against preserving the attention mass over the entire informative set. Retrieval avoids this tradeoff by explicitly selecting which samples enter the context in the first place. Yet in doing so, it uses row similarity as a proxy for the attention that the model would assign to each row. We address the limitations of both curation types with a new model-agnostic mechanism: **Tabular Selective Attention (TSA)**. TSA learns a non-negative score for each test/training-row pair that reshapes the attention distribution into one that favours all informative samples rather than just a subset. Across different architectures, such as TabICLv2 and TabDPT, TSA outperforms existing curation methods on the TabArena benchmark. Leveraging these TSA scores to prune uninformative rows from context, we can speed up its inference while largely preserving accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.