ActVis: Active Maintenance of Structured Visual Evidence for Efficient Visual Chain-of-Thought Reasoning
Abstract
Multimodal chain-of-thought reasoning requires models to acquire and reuse visual evidence across multiple inference steps. Existing training-free Visual CoT methods mainly optimize when and where new visual evidence should be acquired, while previously acquired observations are typically appended to a linear reasoning trajectory without explicit maintenance. As reasoning proceeds, useful evidence may become obscured, whereas redundant, outdated, or conflicting observations continue to occupy the multimodal context. We introduce ActVis, a training-free framework for active visual evidence maintenance. ActVis organizes task-critical observations in a structured and editable Evidence Canvas, where each evidence unit remains grounded in its visual source and can be added, retained, revised, merged, or deleted as reasoning evolves. ActVis further employs an entropy-guided evidence gate to estimate the marginal utility of candidate observations based on predictive uncertainty, retaining informative evidence, filtering redundancy, and revising conflicts. Extensive experiments across multiple multimodal large language models and diverse reasoning tasks demonstrate that ActVis consistently improves reasoning accuracy over linear text–image reasoning and training-free Visual CoT baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.