What Source Calibration Buys Under Real Tabular Shift
Abstract
Real tabular shift breaks the conformal coverage guarantee of tabular foundation models (TFMs). Once a rule reuses the source calibration set, its tail coverage at the same nominal level need not reflect better calibration, since a higher threshold buys coverage with width. We study what the source calibration set is worth for several TFMs on real shifts from US Census, wine and UCI tables. We propose a width-matched control, a reporting protocol that calibrates on the target labels alone at an inflated quantile rank, matches each rule on its mean set size, and summarizes coverage per calibration draw with paired cluster bootstraps over settings and tasks. For source-anchored rules, a bias-variance comparison gives a crossover budget below which the source quantile has lower risk, and it orders the observed crossover across settings, though retrospectively. This supports equal-width comparisons of pooling, prior and gated rules. Experiments show that nominal-level reporting is optimistic for the gated rules, whose gains are mostly width: against the control, a per-draw gain survives at small label budgets and fades as labels grow, pooling rules fall below the control, and the closest prior method sits close to it. Label-free reweighting does not recalibrate real shift under stronger density-ratio estimators or label-shift weights, and the same holds for gradient boosting and logistic regression. Code is available at https://anonymous.4open.science/r/TFM-1-ICLR2027-576E/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.