SIVA-RL: Sensitivity–Invariance Visual Alignment for Multimodal Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards (RLVR) optimizes vision-language models for task-level correctness, but correct answers do not guarantee visual grounding. Existing visual-intervention methods contrast policy behavior on clean and modified images, yet assign supervision by intervention type rather than observed effect. This embeds an assumption we show does not hold. In a paired audit, of clean-correct examples are outcome-drop under one operator and outcome-stable under another. That rate is the disagreement any single per-sample dependence level admits. The required direction belongs to the sample–intervention pair and must be measured, not inferred from operator choice. We propose SIVA-RL, a Sensitivity–Invariance Visual Alignment framework that adds sample-wise, outcome-conditioned supervision through a two-stage training recipe. A frozen audit policy converts the per-sample reward drop into soft routing weights: large-drop pairs receive a margin-bounded sensitivity objective, valid low-drop pairs a bounded consistency objective, and ambiguous pairs are down-weighted. Across seven matched backbone–scale configurations, SIVA-RL raises the nine-benchmark average by to points. These configurations span Qwen2.5-VL at 3B and 7B, Qwen3-VL Thinking at 2B and 8B, and Qwen3.5 at 4B. Gains are larger on the vision-dependent group in five of seven comparisons, with vision-dependent gains of up to points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.