acceptodds
Under review as a conference paper at ICLR 2027

On the Persistent Effects of Lexicality in Large Language Models

Abstract

Representations extracted from large language models (LLMs) play an important role in many downstream applications. However, the structure of these representations is often influenced by lexical overlap rather than semantic content. Our understanding of the relationship between this lexical influence and semantic content, and its implications for downstream tasks, remains limited. In this work, we investigate representations using adversarial semantic stress tests to quantify the effect of lexical overlap relative to semantic content. We find that lexical influence extends across model depth and persists across training regimes and objective functions, including models trained for semantic similarity. Moreover, in autoregressive models we observe an intermediate regime in which semantic discrimination weakens alongside reduced linear accessibility of surface-form information, indicating that reduced lexical accessibility does not imply stronger semantic discrimination. We relate this regime to changes in representation geometry and find that its semantic weakness coincides with a plateau or partial contraction in intrinsic dimension. We further demonstrate the effect of lexical influence on downstream uses of LLMs using summarization and model editing as case studies.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.