Chart-SR: Chart Sensitive Reasoning with Verifiable Rewards in Multimodal LLMs
Abstract
Chart understanding with multimodal LLMs has achieved substantial progress, yet existing methods mainly optimize answer accuracy and lack explicit fine-grained supervision on whether models are sensitive to question-relevant visual evidence. As a result, models may rely on spurious correlations or linguistic priors and struggle to adapt their predictions when critical chart information changes. In this work, we propose a Chart Sensitive Reasoning (Chart-SR) framework that improves visual sensitivity through controlled chart modifications and verifiable reinforcement learning. Chart-SR constructs paired charts consisting of an original chart, a modified chart, and a shared question, where each modification is automatically verified according to its effect on the target answer. Given these pairs, the model learns to identify visual changes, determine whether they are relevant to the question, and update or preserve predictions accordingly. We further introduce a GRPO-based training strategy with multi-component verifiable rewards for answer correctness, change identification, and answer-effect judgment. Experiments demonstrate that Chart-SR improves both general chart reasoning and sensitivity to visual evidence changes, providing a reliable approach for training MLLMs with fine-grained visual supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.