acceptodds
Under review as a conference paper at ICLR 2027

Easy to Call, Hard to Use Effectively: Learning How and When to Use Visual Tools in Small MLLM Agents

Abstract

Multimodal large language models (MLLMs) trained with reinforcement learning with verifiable rewards on tool-using trajectories (Tool-RLVR) have become agents that call visual tools such as zooming and Python code. Small MLLM agents are in growing demand for on-device applications, but turning small MLLMs into such agents remains difficult, as they struggle to use tools effectively. They fail to turn tool results into correct answers and keep calling tools even when the question does not need them, revealing limitations in how and when to use tools. We show that these two limitations hinder Tool-RLVR from training small MLLMs into effective agents, for two reasons. First, Tool-RLVR penalizes a proper tool-call decision whenever a later stage of the trajectory fails, so small MLLM agents decide not to use tools. Second, it ignores whether the tool was necessary for the question, so the call rate moves as a whole rather than following necessity. To address them, we propose Agentic Necessity-Aware Policy Optimization (ANPO), which rewards the decision, execution, and answer stages separately and judges the tool-call decision by the necessity of the tool for each question rather than by the outcome of each trajectory. Extensive experiments on fifteen benchmarks show that ANPO achieves the highest accuracy among Tool-RLVR methods while calling tools mainly where they are necessary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.