Visual Sensitivity Is Not Claim Retractability: Persistence-Aware Credit Assignment for Multimodal Reinforcement Learning
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision–Language Models (LVLMs), and perception-aware methods encourage policies to rely more strongly on visual evidence. Yet stronger visual dependence does not guarantee that individual visual claims are supported by the image: 27.81% of answer-correct base-model trajectories across four benchmarks contain at least one unsupported direct visual claim, which outcome-level RL rewards along with the rest of the trajectory. We introduce a fixed-rollout counterfactual diagnostic that separates Evidence-Function Sensitivity (EFS), how strongly the model responds to an intervened image, from claim persistence, whether it keeps supporting the same claim once its visual evidence is disrupted. The diagnostic reveals Sensitivity–Persistence Decoupling (SPD): under DAPO and VPPO, visual sensitivity increases and aggregate retractability improves, yet unsupported claims become significantly more persistent, whereas GRPO raises sensitivity without this deterioration. We propose Persistence-Aware Credit Gating (PACG), a conservative mechanism that attenuates positive credit for unusually persistent visual claims while leaving negative updates unchanged. It requires no supported/unsupported labels and adds no inference cost. PACG improves unsupported-claim retractability for both optimizers, raises the nine-benchmark average of DAPO from 58.1% to 59.9% and of VPPO from 59.8% to 60.9% on Qwen2.5-VL-7B, transfers to a larger scale and a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual dependence and claim retractability are complementary dimensions of multimodal credit assignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.