The Role of Society of Thought in Vision-Language Reasoning
Abstract
Reasoning models sometimes produce chain-of-thought traces that look less like a single linear solution and more like a dialogue between internal perspectives. Prior work by Kim et a. (2026) has studied this “Society of Thought” (SoT) behavior in text only models, but it remains unclear whether the same phenomenon appears in vision-language models (VLMs), where reasoning must stay grounded in visual evidence. We analyze how this behavior varies across model families, post-training variants, tasks, difficulty, and answer correctness. Society-like reasoning appears across models, but is most visible in models trained to deliberate and in traces for more difficult benchmark items. In contrast to earlier findings in text only models, higher society scores are associated with lower accuracy, including when questions are separated by difficulty. We further distinguish the presence of multiple voices from whether different perspectives communicate and reach a coherent conclusion. Finally, we use society-based rewards to test whether the observed behavior can be changed through reinforcement learning. The rewards successfully alter the reasoning traces and produce modest performance gains without using answer correctness as a training signal. These results show that SoT is present and steerable in VLM reasoning, while its benefit depends on how the different perspectives are coordinated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.