Seeing What Matters: Reward-Anchored Visual Credit Assignment via Self-Distilled Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards improves multimodal reasoning through outcome-level feedback, yet its sequence-level advantage is broadcast across all generated tokens. This obscures which tokens are associated with decisive visual evidence during generation and limits reliable learning from visual evidence. A natural way to address this limitation is to introduce privileged visual evidence and contrast model predictions across different visual conditions, thereby providing finer-grained learning signals for vision-critical tokens. However, the token-level differences induced by these visual conditions do not in themselves constitute reliable credit signals. They may be confounded by factors unrelated to fine-grained visual evidence and may not align with the optimization direction indicated by outcome-level rewards. To this end, we propose **VisCA**, a self-distilled reinforcement learning framework for reward-anchored visual credit assignment. VisCA treats privileged visual differences as candidate evidence for token-level credit assignment rather than directly using them as credit signals. It introduces a matched counterfactual visual condition to assess whether these differences are attributable to fine-grained visual evidence. The verified signal is then calibrated across sibling rollouts. It is subsequently used to modulate the verifier-derived outcome advantage at visually informative tokens, without changing the reward-determined optimization direction. Across six fine-grained visual understanding benchmarks, VisCA improves the average score over the corresponding base models by 7.04 points at 4B scale and 6.21 points at 9B scale. It also maintains strong average performance on held-out tasks that probe broader multimodal capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.