acceptodds
Under review as a conference paper at ICLR 2027

Keep the Image in Mind: A Visual Board for Multimodal Agents

Abstract

Visual reasoning agents must retain useful observations across steps without treating earlier interpretations as substitutes for direct visual evidence. Early visual misinterpretations can persist as faulty premises, leading to incorrect answers despite otherwise coherent reasoning. We introduce a visual reasoning harness centered on the Visual Board, a revisable textual workspace that maintains task-relevant observations and intermediate reasoning across steps. With continued access to the original image, the VLM uses the board to guide targeted visual inspection and revise earlier interpretations as new evidence becomes available. Deterministic history compression bounds textual context, while a comparably sized LLM critic reviews the VLM's interaction trace between rounds and provides feedback on reasoning validity. We evaluate the harness on Euclid 30K, SolidGeo, and the Chart QA subset, covering geometric, spatial, and chart reasoning. Our harness improves accuracy over the corresponding model-only baselines across all evaluated configurations, with gains exceeding 18 percentage points on the Chart QA subset. These results support the effectiveness of our Visual Board-centered harness for improving visual reasoning accuracy. Code is available at https://anonymous.4open.science/r/KeepInMind-1E60/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.