What Is a Visual Thought Worth? When Continuous Visual States Help Vision-Language Models Reason
Abstract
Recent vision-language models augment the visual features of a pretrained backbone with continuous intermediate states intended to support reasoning. These “visual thoughts” can be trained to encode additional visual structure, such as segmentation or depth, or can emerge through iterative latent computation. The goal is to provide the model with a more task-relevant representation of the scene than the backbone features alone may offer. But does this added representation actually work as intended? In particular, can the answer be decoded from it, does the model itself use it, can training make it useful, and does it provide value beyond ordinary image features? We study these questions in two released models, Chain-of-Visual-Thought (CoVT) and Latent Visual Reasoning (LVR). We save each model’s visual states from the original image and progressively weaken or remove the image during answering, while keeping the states fixed. This lets us distinguish three properties that are often conflated: whether an answer can be decoded from a state, whether the released model uses its image-specific content, and whether an answer model explicitly trained to read the states becomes more accurate because of them, relative to an otherwise identical model trained with a non-image-specific mean state. We find that these properties can differ sharply. LVR states become increasingly useful as direct visual evidence disappears, improving counting by 14.2 points and left-or-right accuracy by 12.9 points when the image is removed. CoVT states reveal a different failure mode: the answer is decodable from them even when the answering model initially derives almost no benefit from them, while stronger readout training yields an 8.6-point counting gain without the image. On CLEVR, CoVT states are highly decodable but unused by the released model; after training the answer model to read them, their value grows dramatically as the image degrades, reaching a 71.3-point counting gain with a blank image. Yet at an equal-size interface, ordinary pooled image features outperform the latent states. Because the states are written after the question, part of these gains may reflect an answer already formed, which our runs do not rule out. Thus, adding structured visual states is not sufficient: their value depends on what information is written into them, how that information is exposed to the backbone, and whether subsequent computation learns to use it. Our results motivate latent-reasoning architectures that jointly optimize the representation and its readout, rather than treating the presence of richer intermediate features as evidence that they contribute to reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.