Visual Scaffolding: Toward Native Multimodal Reasoning
Abstract
The success of chain-of-thought (CoT) in complex reasoning can be attributed to a textual scaffolding effect: explicit intermediate results can be examined, reused, and composed to support subsequent reasoning, much as scaffolding supports a building under construction. However, mainstream multimodal LLMs lack an equivalent scaffold for visual distinctions that are difficult to express in concise verbal descriptions. We therefore ask: Can native visual representations become first-class citizens alongside language, providing a visual scaffold for reasoning? We define native multimodal reasoning as reasoning in which native visual representations and language jointly carry reusable intermediate states. As a concrete implementation, the visual scaffolding mechanism we proposed lets the MLLM autonomously initiate a retrieval action at any reasoning step. A lightweight retriever selects native visual features, which are then incorporated directly into the evolving CoT and remain available to subsequent steps. To evaluate this mechanism under controlled visual demands, we introduce NMR-Bench, comprising Burden Counting and Path Tracing. With 88M added parameters (approximately 2%) on Qwen3-VL-4B-Thinking, our model achieves 66.43% and 72.50% accuracy, respectively, surpassing the evaluated GPT-5.5 and Gemini 3.1 Pro baselines. Controlled interventions show that changing selected visual content affects task accuracy and can redirect subsequent visual selection even when intervening text is held fixed. These findings show that native visual states can function as intermediate reasoning material and point toward reusable scaffolds for additional modalities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.