acceptodds
Under review as a conference paper at ICLR 2027

When Does Training Beat Prompting on Tabular Data?

Abstract

Choosing between a prompted language model, a tabular foundation model, and a trained classical model requires understanding where their learning curves exchange advantage and how reproducible that boundary is. We study label-budget crossovers on 18 tabular datasets using 3 splits, shared evaluation rows, and training budgets from 8 to 2,048 labels. Equal-label comparisons are distinguished from comparisons against a fixed 32-shot LLM. We report observed crossover brackets, outcomes outside the measured range, and conditional stability under paired resampling of evaluation rows. For CatBoost classification against GPT-4.1-mini, median bootstrap retention is 57.8% for the exact bracket or censored endpoint, but 99.45% for the broad crossover category: uncertainty in location need not imply uncertainty in whether a crossing is observed. Model choice also changes qualitative outcomes. On the same split, CatBoost achieves lower error than 32-shot GPT-4.1-mini at some measured budget on all 8 regression datasets, but does not overtake 32-shot GPT-5.5 within the evaluated range on 5. A separate retrospective analysis of 126 implementations based on a shared starter protocol examines variation across implementations. These findings motivate reporting crossover location, outcome stability, and model dependence separately when assessing label efficiency in tabular prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.