acceptodds
Under review as a conference paper at ICLR 2027

ActVis: Active Maintenance of Structured Visual Evidence for Efficient Visual Chain-of-Thought Reasoning

Abstract

Multimodal chain-of-thought reasoning requires models to acquire and reuse visual evidence across multiple inference steps. Existing training-free Visual CoT methods mainly optimize when and where new visual evidence should be acquired, while previously acquired observations are typically appended to a linear reasoning trajectory without explicit maintenance. As reasoning proceeds, useful evidence may become obscured, whereas redundant, outdated, or conflicting observations continue to occupy the multimodal context. We introduce ActVis, a training-free framework for active visual evidence maintenance. ActVis organizes task-critical observations in a structured and editable Evidence Canvas, where each evidence unit remains grounded in its visual source and can be added, retained, revised, merged, or deleted as reasoning evolves. ActVis further employs an entropy-guided evidence gate to estimate the marginal utility of candidate observations based on predictive uncertainty, retaining informative evidence, filtering redundancy, and revising conflicts. Extensive experiments across multiple multimodal large language models and diverse reasoning tasks demonstrate that ActVis consistently improves reasoning accuracy over linear text–image reasoning and training-free Visual CoT baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.