Is LLM-Based Tabular Analysis Invariant to Linguistic Realization?
Abstract
Large language models (LLMs) increasingly analyze tabular data through textual representations, allowing them to exploit semantic cues from column names, categorical values, and other natural-language elements. This capability, however, raises a fundamental question: whether LLMs can access and manipulate the same structured information consistently when its linguistic realization changes. We study this problem through a controlled English–Korean setting in which every query is fixed in English and semantically aligned tables are presented in either English or Korean. Our evaluation framework combines paired synthetic and real-world tables, diverse tabular analysis queries, multiple serialization formats, and paired metrics that distinguish aggregate performance differences from instance-level inconsistency. Across multiple LLM families, we find systematic but highly non-uniform sensitivity to linguistic mismatch. English tables generally yield higher accuracy, yet aggregate English–Korean performance gaps can substantially understate behavioral instability: models frequently change correctness across paired instances even when their average accuracy difference is small. Higher overall performance and larger models also do not consistently imply greater cross-lingual stability. Sensitivity varies sharply across query types and is most pronounced in cross-lingual linking, while remaining substantial even for basic column grounding. These results show that aggregate accuracy alone is insufficient to characterize whether LLMs consistently access semantically equivalent tabular information, highlighting the importance of paired cross-lingual evaluation for LLM-based tabular analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.