ContextRL: Addressing Policy Information Bottlenecks with Context-Augmented RL
Abstract
Reinforcement learning with verifiable rewards (RLVR) has become an important route for improving multimodal large language models (MLLMs), using reward signals provided by verifiers as training supervision. Its effectiveness depends on both reliable rewards and positive samples. In multimodal tasks, final-answer agreement can reward flawed reasoning, while difficult queries may yield no correct response within the sampling budget. Our analysis identifies two coupled information bottlenecks that restrict the quality and availability of policy supervision. To address these bottlenecks, we propose ContextRL, a framework that combines context-augmented reward modeling and context-augmented sampling. Context-augmented reward modeling uses full reference solutions to check visual grounding and reasoning, reducing false-positive supervision and generating mistake reports. Context-augmented sampling uses mistake reports to guide further attempts on all-negative groups. Training Qwen3-VL-8B with ContextRL, we achieve average gains of 2.58% and 1.09% over GRPO and DAPO, respectively, across 11 perception and reasoning benchmarks, and surpass Qwen3-VL-32B in perception. Our in-depth analysis shows that contextual information improves reward accuracy and the availability of positive samples, while revealing how falsepositive rewards undermine policy learning. These findings support context augmentation for improving RLVR supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.