A Simple Recipe for Distilling Tabular Foundation Models
Abstract
Tabular foundation models (TFMs) perform strong zero-shot predictions through in-context learning, but their dependence on labeled context makes repeated inference costly. Knowledge distillation offers a way to transfer this predictive strength to lightweight students. However, the context-dependent nature of TFM inference makes the distillation process non-trivial, raising two essential design questions about what constitutes the context and what constitutes the query. We study these questions across two TFMs and both neural and tree-based students, finding a simple and effective recipe for TFM distillation. On TabArena, our distilled students outperform supervised tuned-and-ensembled students by 57–98 Elo points. Moreover, the same distillation recipe improves students on 236–258 of 300 TALENT datasets and reduces median primary error by 4.0–6.4%. Finally, distilled students achieve median inference speedups of 3.0–21.6× over their teachers, making inference substantially cheaper. Code is available at https://anonymous.4open.science/r/TFM_Distillation-ED4A.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.