acceptodds
Under review as a conference paper at ICLR 2027

TabularMath: Diagnosing Computational Extrapolation in Tabular Foundation Models

Abstract

Can a tabular predictor learn a calculation from examples and return its answers when outputs exceed the training range? TabularMath tests this question using 100 GSM8K word problems and 14 AIME competition problems, compiled into reusable computations with 233,472 program-labeled rows. Nine tabular methods and GPT-OSS-120B predict randomly held-out rows or learn from lower-output rows and predict held-out upper-output rows. We measure both regression fit and answer accuracy after rounding to integers. At a 128-row budget under output shift, TabPFN v2.5 achieves median but only 7.8% mean rounded accuracy on GSM8K. The gap occurs within individual computations: 33 families have and accuracy below 20%. Increasing the budget to 2,048 rows improves random-split prediction but leaves shifted accuracy near 8%. GPT-OSS-120B attains 39.2% shifted accuracy at 128 rows, despite negative median from large errors elsewhere. These results make precise answers and controlled error a joint target for computational generalization. We release the executable families and evaluation records.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.