acceptodds
Under review as a conference paper at ICLR 2027

VAST: Visually Anchored Soft Thinking for Multimodal Reasoning

Abstract

Multimodal large language models (MLLMs) increasingly rely on chain-of-thought reasoning, yet autoregressive decoding commits to a single discrete token at each step, limiting exploration of alternative reasoning directions. Soft thinking relaxes this constraint by propagating probability-weighted mixtures of token embeddings, but we find that naively applying it to MLLMs can substantially degrade accuracy due to unnecessary uncertainty. We identify two requirements for effective multimodal soft reasoning: when to activate exploration and how to keep it anchored in visual evidence. We introduce VAST (Visually Anchored Soft Thinking), a training-free framework that addresses both through a position-aware entropy schedule and visual contrastive reweighting. VAST selectively activates soft reasoning during the intermediate reasoning phase and steers the resulting representations toward image-dependent candidates. Across Qwen2.5-VL and LLaVA-NeXT models from 7B to 32B parameters and five multimodal benchmarks, VAST achieves the highest average accuracy on every model, with ablations confirming the contribution of both components. Our results show that effective multimodal soft reasoning requires selective exploration that remains anchored in visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.