acceptodds
Under review as a conference paper at ICLR 2027

How Does Text Help in Tables? A Systematic Evaluation of Text-Rich Tabular Learning

Abstract

Real-world tabular datasets often contain free-form text alongside numerical and categorical features, yet existing methods incorporate such textual information through substantially different interfaces and are typically evaluated under separate protocols. We present a systematic evaluation of text-rich tabular learning that compares four strategies under a unified framework: tabular-only prediction, embedding-augmented learning with frozen text representations, native text-aware modeling, and large language model (LLM) direct prediction across 44 datasets spanning binary classification, multiclass classification, and regression. Through controlled experiments and analyses, we establish three key findings: first, incorporating even basic text representations consistently enhances supervised tabular learners; second, the strongest embedding-augmented model outperforms nearly all native text-aware models; and third, LLM direct prediction underperforms both supervised paradigms due to numerical insensitivity and output compression. Together, these results show that the effectiveness of text-rich tabular prediction depends critically on how textual evidence is represented and integrated with structured supervision. Code and evaluation pipelines are available at https://anonymous.4open.science/r/TextTableBench.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.