CARVEPrep: Model-Aware Data Preparation for Tabular Foundation Model Post-Training
Abstract
Tabular foundation models such as TabPFN are post-trained on task tables, where each record both updates the parameters and serves as context for prediction. Preparing a dirty table thus amounts to allocating a data action to each record. Detecting an error does not show that fixing it helps the model, and action effects need not add, so these decisions should be judged by the trained model rather than by data correctness. We present CARVEPrep, which flags errors from dependency violations and label checks, trains a predictor on self-injected errors, without true values, to choose each record's action, and keeps a prepared table only when repeated short post-training shows that it beats the unchanged one. Feedback from the trained model then reprioritizes records for the next round. We prove that, under the same budget on changed records, removing repair or partial weights can strictly hurt. On six tables with injected errors, CARVEPrep averages 0.8119 accuracy, against 0.7343 for the unchanged data and 0.7853 for the strongest of fourteen baselines, which rewrites 70 times as many cells, and matches or beats every baseline on four of six tables. With data-specific evidence, the same actions carry over to sensor time series, where CARVEPrep ranks first on five of six fault scenarios, and to instruction tuning of Qwen2.5-1.5B, where held-out loss drops from 1.396 to 1.312. Tuning data and post-training hyperparameters together matches or beats tuning either alone on 11 of 12 tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.