Keep the Image in Mind: A Visual Board for Multimodal Agents
Abstract
Visual reasoning agents must retain useful observations across steps without treating earlier interpretations as substitutes for direct visual evidence. Early visual misinterpretations can persist as faulty premises, leading to incorrect answers despite otherwise coherent reasoning. We introduce a visual reasoning harness centered on the Visual Board, a revisable textual workspace that maintains task-relevant observations and intermediate reasoning across steps. With continued access to the original image, the VLM uses the board to guide targeted visual inspection and revise earlier interpretations as new evidence becomes available. Deterministic history compression bounds textual context, while a comparably sized LLM critic reviews the VLM's interaction trace between rounds and provides feedback on reasoning validity. We evaluate the harness on Euclid 30K, SolidGeo, and the Chart QA subset, covering geometric, spatial, and chart reasoning. Our harness improves accuracy over the corresponding model-only baselines across all evaluated configurations, with gains exceeding 18 percentage points on the Chart QA subset. These results support the effectiveness of our Visual Board-centered harness for improving visual reasoning accuracy. Code is available at https://anonymous.4open.science/r/KeepInMind-1E60/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.