acceptodds
Under review as a conference paper at ICLR 2027

TabNSM: Neural Sparse Mixer for Tabular Regression

Abstract

Modern prediction tables increasingly contain hundreds of columns derived from text or large-language-model embeddings, yet only a few matter for any given record. Existing neural models often compare every pair of columns, which becomes costly as tables widen, and models that condition on the training set can exceed practical memory limits. We introduce **TabNSM**, a neural model for regression on these wide tables. Its core *Adaptive Sparse Interaction Module* selects a small, record-specific set of relevant column groups, models their interactions, and combines them with broader table information, so that computation and memory grow linearly rather than quadratically with the number of columns. Across seven highly redundant real-world datasets, TabNSM attains the lowest prediction error on five and a significantly better average rank than each of 41 tuned and pretrained baselines (**rank 1.43 of 42, p ≤ 0.016** against every baseline), with a 21% error reduction on the most interaction-rich table. On the largest tables it trains within a modest GPU-memory budget that does not grow with dataset size, where models that condition on the training set run out of memory or fall back to a small fixed context. Because TabNSM is trained rather than conditioned, its feature attributions can also be validated by removing and retraining on the top-ranked features, which sharply degrades accuracy, unlike attributions from in-context models. An ordinal binning loss (GridLoss) and error-aware sampling (RISE) are secondary training additions. TabNSM is built for the wide, redundant tables that embedding pipelines increasingly produce; tuned trees and in-context models remain ahead on low-dimensional tables and on TabArena.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.