acceptodds
Under review as a conference paper at ICLR 2027

Consistency as an Inductive Bias: Mitigating Cross-View Reliability Gap in MLLMs

Abstract

Inductive biases steer learning toward generalizable solutions by encoding task structure. In this work, we identify a crucial missing bias in MLLMs: cross-view consistency, i.e., semantically invariant views of the same instance should lead to the same answer. Standard reinforcement learning with verifiable rewards (RLVR) objectives do not impose this constraint, but instead assign pointwise rewards to each visual input. Even with data augmentation (DA), transformed views are typically rewarded independently, providing little signal once within-view rewards saturate. We propose ConsistRoll, a simple but effective method that injects cross-view consistency into RLVR training by reusing the group-sampling mechanism of GRPO. Specifically, ConsistRoll places original and semantically invariant transformed views in the same generation group, assigning a joint reward when paired completions are both correct. In this way, ConsistRoll turns consistency into an online credit-assignment signal, without extra annotations and overhead. Theoretically, we prove that a reward computed from finitely many paired responses recovers the gradient of this long-run reliability target on average, and quantify how same-group centering modifies the signal. Comprehensive 14 benchmarks across math, general-purpose, and hallucination domains confirm that ConsistRoll achieves robust improvements in cross-view stability and multimodal reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.