When VLMs Rely Less on Their Reasoning: Adaptive KL Regularization for RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) has shown promise for improving reasoning in vision-language models (VLMs). Productive exploration is important for continued performance gains, motivating efforts to relax reference-policy constraints. However, removing reference-policy KL regularization in VLM RLVR can degrade reasoning–answer consistency even when final-answer accuracy remains comparable. Our analyses suggest that RLVR reduces VLMs' reliance on generated reasoning relative to visual inputs during answer prediction. To address this issue, we propose PEAK (Position- and Entropy-Aware KL Anchoring), a token-adaptive KL regularization method that uses the pre-RLVR model as a reference to help preserve reliance on generated reasoning while encouraging productive exploration. PEAK is guided by two observations: (1) reasoning–answer inconsistent responses exhibit larger actor–reference entropy gaps, particularly without KL regularization; so PEAK assigns stronger regularization to tokens with larger entropy gaps; (2) large positive entropy gaps are concentrated in a subset of tokens, while token entropy is elevated at later response positions; PEAK adapts its regularization strength to both token-level entropy gaps and response position. Across seven multimodal reasoning benchmarks, PEAK outperforms the evaluated RLVR methods in average final-answer accuracy. It also improves reasoning–answer consistency and Pass@ over both uniform-KL and no-KL baselines, with gains of 21.04 and 7.90 percentage points in consistency and Pass@, respectively, over standard DAPO (without KL regularization). These joint improvements support more productive exploration while better preserving reasoning behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.