Improving Generalization Robustness of Multimodal RLVR
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models (MLLMs) more accurate, but the gains are brittle: simply rewording the instruction or changing the answer template can erase them, which challenges reliable deployment in high-stakes scenarios like medical VQA. The base model barely reacts to such changes, so the brittleness is created by RL post-training itself. We trace it to two issues of the standard RLVR objective. First, the binary verifier conflates format with content, so the reward signal cannot tell a wrong answer apart from a misformatted one. Second, the training prompts cover only a thin slice of the semantically equivalent prompts met at deployment, so a policy that is optimal on the training template can behave differently on unseen ones. Both failures call for an objective over the full class of semantically equivalent prompts, and we identify two conditions that help achieve it: separating format from semantics in the reward, and enforcing answer invariance across equivalent prompts. We therefore propose Prompt-Invariant RLVR (PIRL), consisting of a dynamic trinary reward and a consistency regularizer driven by an embedding-space adversary. Under template stress testing, PIRL's mean accuracy drops by only –%, whereas GRPO drops –. On dynamic evaluation, PIRL also achieves the smallest drop among trained methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.