acceptodds
Under review as a conference paper at ICLR 2027

Localized GRPO: Utilizing Dense Reward Signal From LLM-Judge Error Localization

Abstract

Reinforcement learning for LLMs is shifting from PPO toward critic-free, group-relative methods such as GRPO, which remove the value network that PPO uses to form token-level advantages. Vanilla GRPO instead forms one group-relative advantage per rollout and broadcasts it to every token, so the algorithm has no explicit mechanism for assigning reward credit within a response. On verifiable reasoning tasks, recent methods address this limitation using intrinsic proxies such as token entropy, likelihoods, or Monte-Carlo rollout estimates. We study a complementary setting: generation tasks where external LLM judges can identify localized quality signals in a response, such as error spans in machine translation (MT) and supported or hallucinated spans in long-form factuality. We convert these explicit fine-grained span judgments into per-token rewards, and introduce L-GRPO, a localized GRPO advantage estimator that uses those rewards while preserving GRPO's sequence-level group-relative signal. On document-level MT across ten language pairs, L-GRPO reduces LLM-judged error density by 7.2% relative to GRPO. On long-form factuality, it improves factual precision across four benchmarks by 9.6% relative. Both gains are statistically significant. L-GRPO also outperforms two strong dense credit baselines on both tasks: PPO+GAE, which uses a learned critic, and SDPO, which adds feedback-conditioned self-distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.