acceptodds
Under review as a conference paper at ICLR 2027

Neither Object nor Text: A Mechanistic Analysis of Scene Text in LVLMs

Abstract

Large vision-language models (LVLMs) are increasingly required to recognize scene text embedded in images. Scene text is spatially localized like visual objects, yet conveys linguistic content like textual input. This dual nature motivates us to investigate how scene text is internally represented and processed within LVLMs. Through empirical analyses of LVLM computation, we identify processing differences between scene text, visual objects, and directly supplied text. On identical images, scene text questions induce stronger attention concentration in the corresponding regions than object questions. Beyond this attention pattern, hidden-state analysis shows that scene text information is already detectable in visual representations before generation begins. Attention knockout further shows that scene text recognition depends more strongly than object recognition on access to image tokens in later decoder layers. Scene text also differs from directly supplied text in its layerwise attention patterns. Together, these findings suggest that scene text may function as a distinct representational modality within LVLMs: neither fully object-like nor purely language-like. These findings motivate further investigation into how scene text representations guide answer generation and how this understanding can improve recognition and response behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.