acceptodds
Under review as a conference paper at ICLR 2027

What Do VLM Language States Remember? Low-Rank Recovery of Instance-Level Visual Evidence

Abstract

Does a multimodal language model retain, in its contextualized language states, evidence pointing to the specific visual instance a referring expression denotes: the dog on the left, not merely dog? We probe this with a capacity-limited linear decoder, asking whether language states can linearly recover the corresponding spatially localized visual evidence from frozen vision representations. Across Qwen3-VL, LLaVA-OneVision, and Llama-3.2-Vision on RefCOCO, a low-rank linear mapping recovers the instance-level correspondence, reaching 88.4%, 88.6%, and 71.7% Pointing, with most gains saturating near rank 32; rank sweeps and cross-modal geometry localize this structure to a few shared dimensions, while native-attention, CLIP, and statistical-projection baselines perform substantially worse. CKA/CCA, semantic mismatch, anchored concept swapping, and unseen-category transfer further show that recovery follows cross-modal structure already present in frozen representations rather than arbitrary bbox associations. Spatially matched suppression and retention show that the recovered evidence has a location-specific causal effect on outputs in Qwen and LLaVA, distinguishing decodable from functionally used information. Token-level analyses show instance specificity emerges only as the expression is contextualized (bike sharpening into bike on the far left), and across architectures the location and strength of recoverability depend on how visual representations are injected into the language stream. Thus, the models retain, in their language states, a compact low-rank structure of linearly recoverable instance-level visual evidence, beyond what native attention reveals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.