acceptodds
Under review as a conference paper at ICLR 2027

LATTE: Bridging LLMs and Tabular Foundation Models for Tabular Learning with Strings

Abstract

Real-world tables often mix numerical features with free-form strings. Table foundation models (TFMs) excel on numerical tables but require strings to be encoded numerically, while large language models (LLMs) hold the world knowledge those strings refer to. Most existing bridges use the LLM as a plug-in encoder that takes an isolated cell in, and gives a mean-pooled last-layer vector out. We show that this *interface* between LLM and table matters as much as the LLM itself. Read one cell at a time, a larger LLM does not give better features: on average, Qwen3-0.6B even beats Qwen3-8B. With our interface, built from three ingredients—*row serialization* (giving every cell its row-level context), *intermediate-layer extraction* (avoiding layers specialized for next-token prediction) and *columnar pooling* (removing mean pooling's length bias)—the same LLMs improve steadily with size. With *by-pass features* routing numbers and exact cell identity past the LLM, these ingredients form LATTE (LLM-Augmented Table Text Embeddings). Paired with TFMs, LATTE sets a new state of the art for tables with strings. Across 100 diverse tables, it beats the strongest baseline on 89% of splits, with gains across 7 LLMs and 10 tabular learners, and ranks first on the 108-dataset STRABLE benchmark.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.