Rethinking Visual Thoughts: From Inert Context to Privileged Supervision
Abstract
Interleaved visual chain-of-thought (CoT) enables multimodal models to generate intermediate images during reasoning, under the premise that visual states capture what text poorly describes. We investigate a fundamental question: do these models actually use the images they generate? Through a targeted intervention framework, we reveal that self-generated visual thoughts remain largely question-invariant, and replacing them with noise or the original input causes little performance drop. Instead, subsequent reasoning is overwhelmingly dominated by language, treating visual tokens as inert context. Building on this insight, we propose treating visual thoughts as privileged supervision during training to transfer visual knowledge directly into the language pathway. We explore two distinct strategies: explicitly verbalizing visual states into text (Verbalized CoT, V-CoT) and implicitly aligning the hidden states of the student's own reasoning with a teacher that reads the visual thoughts (Visual-State Distillation, VSD). Extensive evaluations across diverse model backbones show that both strategies yield highly effective text-only reasoners, outperforming standard interleaved baselines on average in every recipe across varied visual reasoning domains while substantially improving overall inference efficiency at test time. All code, data, and model weights will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.