acceptodds
Under review as a conference paper at ICLR 2027

RECAST: RESIDUAL-STATE REPRESENTATION LEARNING WITH FOUNDATION MODELS FOR LIGHTWEIGHT TABULAR PREDICTION

Abstract

Even well-trained tabular models can exhibit systematic error patterns in particular regions of the feature space. We propose RECAST, which learns residual-state representations to improve an existing model, rather than directly relearning its residuals from the original features alone. At each correction stage, RECAST partitions out-of-fold XGBoost errors into three states. An LLM uses feature semantics and aggregated error statistics to select feature subsets and propose a focus condition that guides in-context example selection, while a tabular foundation model (TFM) estimates residual-state representations for new inputs. A teacher augments the original features with these representations to predict residual corrections, and the residual states and context are updated after each stage. The teacher’s out-of-fold predictions are then blended with ground-truth labels to train a single XGBoost student using only the original features. Foundation models are used only during training; inference requires only the student and the original inputs. Across nine regression and nine binary classification datasets from OpenML, RECAST achieves the best average rank among ten compared models: 4.11 for regression, 3.00 for classification, and 3.56 overall. It improves the baseline XGBoost’s three-seed mean performance on all 18 datasets. Controlled comparisons further show better task-level mean performance than residual relearning without state representations and direct distillation of the TFM’s original-target predictions. These results support residual-state estimation as a training-time use of foundation models for improving an existing tabular predictor while retaining lightweight inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.