Amortizing Label Shift for In-Context Tabular Prediction
Abstract
Prior Fitted Networks have emerged as a disruptive paradigm enabling the pre-training of Tabular Foundation Models (TFMs) entirely on synthetic data. While these models have proven successful across a broad variety of real-world predictive tasks, they still make strong assumptions about the train and test distributions being fixed. This assumption is often unrealistic, as many applications exhibit substantial distribution shifts. In particular, a drawback of current TFMs for real-world deployment is their vulnerability to label shift, where class proportions change between train and test time. In this work, we introduce ShiftICL, the first Tabular Foundation Model that explicitly amortizes label-shift adaptation during pre-training. ShiftICL is trained on synthetic tasks spanning a broad range of label distributions and shift severities, enabling it to adapt its predictive distribution to previously unseen label shifts without retraining or access to test labels. Across a large collection of 205 OpenML tabular datasets under controlled label shift, ShiftICL substantially improves predictive accuracy and calibration over existing state-of-the-art methods tailored for distribution drifts, with gains becoming increasingly pronounced as the shift grows.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.