acceptodds
Under review as a conference paper at ICLR 2027

DIA-GRPO: Guiding Policy Optimization with Cross-Prompt Diagnosis

Abstract

Group Relative Policy Optimization (GRPO) improves reasoning in large language models through efficient, group-based reinforcement learning. However, its prompt-local advantage construction leaves a key source of learning information untapped: how successes and failures across prompts jointly reveal the model's evolving capabilities and inform how individual outcomes should contribute to its overall improvement. To address this limitation, we introduce Diagnosis-Informed Advantage for GRPO (DIA-GRPO), which translates diagnosis of the LLM's evolving capabilities into training feedback. By jointly modeling training-sample attributes and the LLM's evolving capabilities, DIA-GRPO translates cross-prompt observations and a population improvement objective into response-level advantages. Specifically, DIA-GRPO evaluates each outcome against the model's expected performance and uses the overall improvement objective to determine how that outcome contributes to policy updates. DIA-GRPO is lightweight and plug-and-play: online diagnosis operates in a low-dimensional capability space, and the resulting advantages are mixed with native GRPO advantages within the existing policy optimization objective, requiring no additional verifier or policy-scale critic. Theoretically, under explicit local conditions, we show that the expected unclipped diagnostic update recovers the policy gradient of a population log-success objective. Experiments on mathematical reasoning show that DIA-GRPO improves overall mean correctness by 2.96 and 2.05 percentage points on Qwen3-1.7B and Qwen3-8B, respectively. The diagnostic module also delivers consistent gains across different policy optimization methods with less than 2% additional training time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.