When Does Chain-of-Thought Hurt Spatial Answers? Length, Retrieval, and Content Controls in Vision-Language Models
Abstract
Chain-of-thought (CoT) prompting often lowers the spatial accuracy of vision-language models (VLMs). In a decoder-only VLM, reasoning text cannot change the image or question states, so it can act only through how the answer position retrieves them or through what the reasoning says. We separate these routes on 2850 spatial items (CV-Bench and two synthetic tasks) in 5 open VLMs of at most 8B parameters with length, turn-structure, re-injection, regeneration, and content controls, reading every answer as one forced-choice token. CoT changes accuracy by -5.7 points on average, whereas length-matched neutral filler in a separate turn changes it by -0.6 and 1024 filler tokens by -2.4; restating the question or re-supplying the image after the reasoning recovers at most +0.8 and +1.0 points. The answer agrees with the option the reasoning states, and in four models pooled CoT accuracy is within 1.7 points of p̄, the Direct distribution's mean probability of the correct option: CoT scores like a sample from that distribution rather than its argmax. Turn-matched filler keeps Direct's argmax step at a correct-option probability of 0.5 while CoT flattens it, and the same probabilities predict four-sample majority-vote accuracy within 1.4 points. The CoT effect thus decomposes into the argmax–expectation gap G = Direct − p̄, which CoT gives up, and the gain E = CoT − p̄ beyond the expectation. In nine larger or closed VLMs queried through an API (reasoning disabled where the API allows it; GPT-5 models still reason in part of their direct calls), G is at most +2.9 points, E reaches +11.2, and CoT changes accuracy by -0.7 to +11.7 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.