Rethinking Group Accuracy: Latent-Reliability Inference for Prompt Solvability in RLVR
Abstract
In reinforcement learning with verifiable rewards (RLVR), group relative policy optimization (GRPO) constructs relative advantages by centering response rewards with the empirical group-average reward, a finite-sample estimate of prompt solvability under the current LLM. However, empirical group accuracy depends only on the success count and treats equally rewarded responses as equally informative. Under heterogeneous response reliability, this can misestimate prompt solvability and miscalibrate relative advantages. Consistent with this concern, we empirically observe substantial reliability variation among prompts with identical group accuracy, revealing that the same success count can mask markedly different response-level reliability patterns. In this work, we formulate prompt-solvability estimation as a latent-variable inference problem and introduce Reliability-Aware Latent Solvability Inference (RLSI). Our approach introduces a latent response-level success state and relates each verifier outcome to this state through its response-specific reliability. Under this latent-variable model, we derive an efficient expectation–maximization (EM) procedure to estimate prompt solvability and use the resulting estimate for advantage construction. Theoretically, we show why reward-only aggregation can fail to recover prompt solvability, while response-level reward–reliability pairings retain information relevant to latent-state inference; we further characterize the resulting reliability-dependent corrections and recover the GRPO group mean in the perfect-reliability limit. Across six mathematical reasoning benchmarks and two model scales, RLSI achieves the strongest average performance among the evaluated methods, outperforming GRPO by 5.2 and 4.8 points on Qwen3-1.7B-Base and Qwen3-4B, respectively. Additional comparisons and ablations further support the effectiveness of latent solvability inference and the value of response-level reliability heterogeneity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.