acceptodds
Under review as a conference paper at ICLR 2027

When VLMs Rely Less on Their Reasoning: Adaptive KL Regularization for RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) has shown promise for improving reasoning in vision-language models (VLMs). Productive exploration is important for continued performance gains, motivating efforts to relax reference-policy constraints. However, removing reference-policy KL regularization in VLM RLVR can degrade reasoning–answer consistency even when final-answer accuracy remains comparable. Our analyses suggest that RLVR reduces VLMs' reliance on generated reasoning relative to visual inputs during answer prediction. To address this issue, we propose PEAK (Position- and Entropy-Aware KL Anchoring), a token-adaptive KL regularization method that uses the pre-RLVR model as a reference to help preserve reliance on generated reasoning while encouraging productive exploration. PEAK is guided by two observations: (1) reasoning–answer inconsistent responses exhibit larger actor–reference entropy gaps, particularly without KL regularization; so PEAK assigns stronger regularization to tokens with larger entropy gaps; (2) large positive entropy gaps are concentrated in a subset of tokens, while token entropy is elevated at later response positions; PEAK adapts its regularization strength to both token-level entropy gaps and response position. Across seven multimodal reasoning benchmarks, PEAK outperforms the evaluated RLVR methods in average final-answer accuracy. It also improves reasoning–answer consistency and Pass@ over both uniform-KL and no-KL baselines, with gains of 21.04 and 7.90 percentage points in consistency and Pass@, respectively, over standard DAPO (without KL regularization). These joint improvements support more productive exploration while better preserving reasoning behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.