acceptodds
Under review as a conference paper at ICLR 2027

ContextRL: Addressing Policy Information Bottlenecks with Context-Augmented RL

Abstract

Reinforcement learning with verifiable rewards (RLVR) has become an important route for improving multimodal large language models (MLLMs), using reward signals provided by verifiers as training supervision. Its effectiveness depends on both reliable rewards and positive samples. In multimodal tasks, final-answer agreement can reward flawed reasoning, while difficult queries may yield no correct response within the sampling budget. Our analysis identifies two coupled information bottlenecks that restrict the quality and availability of policy supervision. To address these bottlenecks, we propose ContextRL, a framework that combines context-augmented reward modeling and context-augmented sampling. Context-augmented reward modeling uses full reference solutions to check visual grounding and reasoning, reducing false-positive supervision and generating mistake reports. Context-augmented sampling uses mistake reports to guide further attempts on all-negative groups. Training Qwen3-VL-8B with ContextRL, we achieve average gains of 2.58% and 1.09% over GRPO and DAPO, respectively, across 11 perception and reasoning benchmarks, and surpass Qwen3-VL-32B in perception. Our in-depth analysis shows that contextual information improves reward accuracy and the availability of positive samples, while revealing how falsepositive rewards undermine policy learning. These findings support context augmentation for improving RLVR supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.