INCEPT: Infusing Novel Concepts for Explaining Pretrained Tabular models
Abstract
The Platonic Representation Hypothesis suggests that sufficiently capable models converge to similar internal representations regardless of architecture or training. Using embeddings from the TabArena benchmark to assess geometric convergence of tabular foundation models, we find transformer-based tabular foundation models (TFMs) cluster tightly, with hypernetworks and LLM-based models sitting apart. To explain the predictive gap between two TFMs, we train a sparse autoencoder (SAE) per model and identify their unmatched concepts. Two interventions are tested: ablation zeros unmatched concepts out of the stronger model's embedding; transfer maps them into the weaker model's embedding. Across 15 model pairs and 51 TabArena datasets, ablation closes 93% of the stronger model's per-row advantage, and transfer recovers 90% of the gap. The transferred concepts lie roughly half off the weaker model's existing manifold, revealing structure the recipient does not use but can absorb.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.