acceptodds
Under review as a conference paper at ICLR 2027

BiVER: Bidirectional Visual Evidence Rewards for Multi-Tool VLM Reasoning

Abstract

Vision Language Models (VLMs) can actively manipulate images through tool calls during reasoning, yet common reinforcement learning objectives provide little direct feedback on the images these tools produce. An irrelevant tool output thus inherits a positive reward whenever the answer is recoverable without it, while an informative one receives no credit when the final answer is wrong. We show that a frozen VLM can assess returned visual evidence along two complementary directions: a *forward* answer-utility score measures support for the reference answer, while a *backward* question-alignment score measures relevance to the question without using the answer. Building on this finding, we propose **BiVER** (**Bi**directional **V**isual **E**vidence **R**eward) for multi-tool VLM reasoning. It scores intermediate tool outputs in both directions and incorporates their trajectory-level aggregate into GRPO. Because BiVER evaluates the image produced by an executed action rather than its name or arguments, it provides a shared reward criterion for heterogeneous image-producing tools without supervised tool-use trajectories or training a dedicated reward model. Trained from Qwen2.5-VL-7B-Instruct with six tools, BiVER improves its initialization by an average of **11.9** points across high-resolution benchmarks and **5.0** points across document, chart, and table benchmarks, achieving the highest average accuracy among the compared models for each benchmark suite.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.