acceptodds
Under review as a conference paper at ICLR 2027

Reward Hacking in Multimodal Reinforcement Learning via Linguistic Priors

Abstract

Reinforcement learning (RL) has become a powerful post-training paradigm for multimodal large language models (MLLMs), which are increasingly deployed in safety-critical domains. However, we reveal a critical yet overlooked phenomenon: models optimized via RL may exploit linguistic priors encoded in their language backbone to predict answers without truly processing image content. We term this behavior Visually-bypassed Shortcut Learning (VSL). Unlike reward hacking in text-only models, VSL is invisible to existing monitoring, as both pathways yield superficially identical reasoning traces. We show that this shortcut is systematically favored during RL optimization due to its lower computational cost, reduced reward variance, and the rich statistical structures of the language backbone, causing RL training to actively degrade visual grounding. Using image-perturbation sensitivity metrics, we verify the pervasiveness of VSL across multiple MLLMs and task types in clean training environments, and find its severity strongly correlates with task-specific linguistic prior strength. These findings expose a silent trade-off: benchmark scores keep improving while visual grounding steadily erodes, yet no existing signal indicates when the degradation becomes unacceptable. We therefore propose Visual Information Evaluation (VIE), a low-cost detector requiring only forward passes and no training modification, which exploits the insight that shortcut-reliant models need less visual information to maintain high rewards. Experiments show that VIE reliably detects visual reward hacking under both in-distribution and out-of-distribution settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.