acceptodds
Under review as a conference paper at ICLR 2027

DivR: Grounding LLM Reinforcement Learning in Rollout Behavioral Divergence

Abstract

LLM-as-a-judge has become a practical approach for providing rewards in reinforcement learning on open-ended tasks, where correctness is difficult to verify automatically. A common approach is to let a judge LLM observe current rollouts and generate quality rubrics for scoring them. We argue that rubrics designed for evaluation should not be directly reused for training: evaluation rubrics only need to judge the absolute quality of an individual response, whereas training rewards must also distinguish the relative quality of rollouts within the same group. We propose Divergence-Point Reward (DivR), which identifies quality-relevant behavioral divergences among current rollouts and scores responses relative to their peers. Across HealthBench and WritingBench, DivR outperforms both judge-generated and dataset-specific rubric rewards, improving over GenR by 11.6–21.8% under independent HealthBench judges and achieving the strongest ranking across judges on WritingBench. Mechanistic analyses show that divergence-grounded reward dimensions are more reliable and preserve informative group-relative signals as the policy improves. Our code is available at https://anonymous.4open.science/r/divr_public-FA48.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.