acceptodds
Under review as a conference paper at ICLR 2027

ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

Abstract

Multimodal scientific claim verification (MSCV) checks claims against text, tables, and figures of scientific papers. It requires extracting multiple pieces of localized evidence and combining them through multi-step reasoning; vision-language models (VLMs) struggle with both. Generic crop-and-zoom tools only return pixels, and prompting VLMs to use tools can even hurt accuracy. We present \ours, which teaches VLMs when and how to acquire claim-relevant evidence. It pairs SciEviKit, a type-aware tool suite returning structured evidence (table rows or columns, chart values, zoomed regions), with two-stage training: On-Policy Self-Distillation (OPSD) bootstraps tool use from reference tool calls alone, and Group Relative Policy Optimization (GRPO) optimizes a reward coupling accuracy with tool-call validity and efficiency. We evaluate on SciVer and MuSciClaims, spanning diverse evidence modalities, domains, and reasoning structures, with three backbones (Qwen, InternVL, Gemma) and five baselines covering non-tool reasoning, training-free agents, and RL-based tool use. \ours improves over non-tool chain-of-thought by 8.16 points on SciVer and 7.92 on MuSciClaims, and over the strongest baselines by 3.70 and 2.23 points. Our analyses further show better evidence acquisition, tool selection, and evidence-based reasoning, and ablations confirm the value of OPSD warm-up.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.