acceptodds
Under review as a conference paper at ICLR 2027

ProVIC: Provenance-Aware Credit Assignment for Visual Tool Reasoning

Abstract

Group-based reinforcement learning (RL) trains visual-tool agents from final-answer rewards, but typically assigns the same trajectory-level credit to tool calls with different roles in evidence acquisition. Local group-based methods distinguish these calls by comparing returns under matched intermediate conditions. For visual-tool calls, however, neither ordered histories nor returned values reliably identify equivalent evidence. In particular, different acquisition orders can yield equivalent prior evidence, while identical values can arise from different image regions or calculations. To address this matching problem, we introduce ProVIC (Provenance-aware Visual Interaction Credit), which converts final-answer rewards into call-specific learning signals for visual-tool segments. By using a persistent evidence graph, ProVIC matches equivalent prior evidence across acquisition orders while preserving observation provenance. It then uses matched final-answer rewards to give each tool segment a call-specific advantage. Experiments with Qwen3-VL-8B-Instruct on five visual reasoning benchmarks show that ProVIC achieves superior performance over all the compared methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.