acceptodds
Under review as a conference paper at ICLR 2027

A Controlled Study of Scaling Training Data in Tabular Machine Learning

Abstract

Despite increased attention to the field of tabular machine learning, the niche of large-scale (1M+) data remains under-explored. In this work we study training-data scaling on eight datasets with up to ten million examples, evaluating tabular ML methods under realistic distribution shifts with in-distribution (ID) controls. First, we find that GBDTs and MLP-based models generally improve with additional data and exhibit broadly similar scaling trends. Tabular foundation models vary substantially across pretrained versions, with recent models benefiting more consistently from larger contexts, in some cases up to one million examples under distribution shift. Our analysis of MLP-based models shows that, as training data scales, ID and out-of-distribution (OOD) performance can diverge: for MLP-based models selected using ID validation, additional data can improve ID performance while degrading OOD performance. We show that target-aligned validation mitigates this effect through tuning and early stopping. In contrast, tabular foundation models lack a commonly adopted mechanism for using validation as a separate feedback signal, and we demonstrate that in the presence of strong distribution shift a common strategy of adding validation examples to the context does not prevent OOD degradation as training data grows.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.