acceptodds
Under review as a conference paper at ICLR 2027

Disentangled Value Estimation: Rethinking Scalar Critics in Rubric-Based RL

Abstract

Value estimation is crucial for many reinforcement learning algorithms used to optimize large language models (LLMs). In critic-based methods such as Proximal Policy Optimization (PPO), an additional critic model predicts expected future outcomes for each generated token based on intermediate generation states, and this value is used to compute the advantages that guide policy optimization. Such value estimation is relatively well defined when the reward specification is fixed and consistent, but becomes more challenging in open-ended tasks where evaluation criteria vary across queries. We study this problem in RL with rubric rewards, where each query is associated with multiple query-specific criteria whose scores are collapsed into a scalar reward. A standard critic observes only the query and response prefix and is supervised by this aggregated reward, requiring it to implicitly recover the underlying reward specification before estimating values. To address this, we propose Disentangled Value Estimation (DVE), which uses a shared critic conditioned on one rubric at a time to estimate rubric-specific values, computes token-level advantages for each rubric, and aggregates them using the original rubric weights. Across five benchmarks spanning instruction following, chat, and scientific question answering, DVE outperforms the vanilla PPO and GRPO in 13 of 15 model-benchmark combinations. Further analysis shows that DVE achieves lower value prediction error, higher explained variance, stronger response-level ranking consistency, and more targeted token-level credit assignment. Code is available at Anonymous Repo.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.