acceptodds
Under review as a conference paper at ICLR 2027

Do You Even Need More Data? Rethinking LLM Deployment for Label-Scarce Tabular Classification with TabLens

Abstract

Every tabular classification task requires sufficient labeled data. When labels are scarce, practitioners can collect more examples, encode domain expertise into features, or accept lower predictive performance. Large language models offer a potential alternative by leveraging knowledge acquired during pretraining. This possibility raises three distinct questions: (i) can zero-shot inference reduce the need for labeled data; (ii) can a richer prompt reduce the need for collecting more examples; (iii) and can encoded domain knowledge compensate for missing data? To answer each question under explicit, testable conditions, we introduce TabLens, a benchmark of 91 real-world and synthetic datasets across eight domains, together with eight industrial datasets containing free-text columns. First, below a complexity threshold that can be estimated from a dataset's intrinsic dimensionality before running an LLM, zero-shot inference matches the accuracy that would otherwise require up to approximately 16 labeled examples to collect. This holds whether the model generates a written answer or, more cheaply, scores candidate labels directly from their log-probabilities – the latter loses almost no accuracy except for the smallest models. Above the threshold, this advantage disappears. However, on datasets with free-text columns, zero-shot LLMs are even better than 16-shot TabPFN-3 Plus, and no longer require manual feature engineering. Second, a richer prompt reduces the number of required demonstrations only when it includes both the real column names and a plain-language task description. Providing only one of these sources of context can perform worse than providing neither, particularly at higher shot counts. Third, on the LLM-synthetic datasets, explicit if–then expert decision rules improve performance only when the prompt provides sufficient task context for the model to interpret them. Even then, they only substitute for up to about 16 real examples, after which collecting more data is more effective. Together, these results establish testable conditions under which LLMs can reduce data-collection and feature-engineering requirements for tabular classification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.