acceptodds
Under review as a conference paper at ICLR 2027

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

Abstract

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks. In this work, we rethink agentic visual reasoning through two key dimensions of tool use: **Mode Adaptiveness** and **Tool Effect**. Mode Adaptiveness characterizes whether an MLLM can recognize when tools are truly necessary and invoke them accordingly, thereby avoiding unnecessary computational overhead while improving performance on challenging problems that require tool assistance. Tool Effect characterizes the actual impact of tool use: tools should extend the model's capabilities on problems unsolvable through tool-free reasoning, while avoiding additional errors on problems that the model can already solve without tools. We conduct a comprehensive analysis to quantify these two properties and empirically reveal that existing agentic visual reasoning models exhibit limited Mode Adaptiveness, while the gains produced by tool use on hard examples are largely offset by the harm introduced on easy examples that the models can already solve. Motivated by these observations, we propose **Beacon**, a novel agentic visual reasoning model trained with supervised fine-tuning (SFT) and reinforcement learning (RL). At the core of Beacon are the *Necessity-Aware Adaptive Reward* and the *Hint-Guided Capability Expansion* mechanism in the RL stage. *Necessity-Aware Adaptive Reward* encourages tool-free solutions when they succeed while preserving full reward for successful tool use when tool-free rollouts fail. Hint-Guided Capability Expansion uses verified, answer-free expert hints to recover learning signals from all-wrong rollout groups, aiming at extending the tool-use capability on the hardest problems. Across 13 benchmarks, Beacon achieves the highest average score among the evaluated open-source models and ranks first on 11 benchmarks. On five diagnostic benchmarks, it improves the average tool-available accuracy over its tool-free accuracy by 1.96 points and achieves the largest tool-gain minus tool-harm score (+3.14 points). These results show Beacon's advanced performance, Mode Adaptiveness, and the net benefit of tool use.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.