OpenVLHarness: Real World Visual Agentic Tool-Use with Multimodal Memory
Abstract
Multimodal large language models (MLLMs) have made rapid progress on visual reasoning, yet many tasks require evidence that is difficult to obtain through internal reasoning alone—our pilot study on 12 benchmarks shows that simply increasing the inference budget yields only modest gains—such tasks demand grounded perception, external knowledge, and numerical measurements obtained through heterogeneous tools, whose outputs must be made interpretable to the model and retained across reasoning steps. We introduce OpenVLHarness, a multimodal reasoning harness organized around a capability layer and persistent multimodal memory. The capability layer combines task-level tool interfaces with evidence rendering, providing consistent access to heterogeneous backends and converting their outputs into model-readable observations. Persistent multimodal memory retains the corresponding images, structured data, and web references as addressable artifacts throughout each query. A coding interface executes generated programs over stored data, enabling computations that combine results from multiple tools. Together, these mechanisms support an iterative process of acquiring, inspecting, and composing multimodal evidence. Evaluated on 23 benchmarks across perception, spatial reasoning, visual search, and multimodal question answering, OpenVLHarness consistently surpasses strong MLLM backbones and outperforms existing visual tool-use baselines across six open and proprietary backbones of different scales, including +11.2 with Qwen3-VL-8B, +9.2 with Kimi K3, and +11.2 with GPT-6 Sol. In addition, our controlled study on 12 benchmarks shows that even though OpenVLHarness incurs additional cost from multi-step tool use, giving the base GPT-6 Luna an even larger reasoning budget (1.27x the cost of OpenVLHarness) improves the average only from 49.6 to 51.5, compared with 62.9 for OpenVLHarness. These results suggest that additional reasoning alone does not substitute for acquiring and composing task-relevant multimodal evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.