acceptodds
Under review as a conference paper at ICLR 2027

Does GRPO Acquire or Redistribute? Diagnosing Gains Beyond Observed Base Support

Abstract

Improved answer accuracy after reinforcement learning does not by itself distinguish acquired capability from a redistribution of probability over solutions already accessible to a base model. We examine this distinction through an artifact audit of a Qwen3-8B mathematical-reasoning study. Historical score vectors identify ten problems with zero base credits and 75 trained credits at 128 samples per problem. A subsequent evaluation preserves 20,480 generations from these problems. We exactly reproduce its 19 base and 12 trained automatic credits, but every generation reaches the 1024-token limit. Response inspection exposes coefficient matches, scalar/root-pair normalization collisions, and an uncredited prefix that derives the correct value. Consequently, neither the automatic counts nor an audit restricted to credited responses yields a reliable estimate of answer correctness. A separate forked-update evaluation records differences of one credit in each of two probe groups, but its implementation is not verified standard GRPO, and adapter identity and train–probe separation remain insufficiently established for causal interpretation. We provide deterministic replay, traceable counterexamples, and an evidence checklist for beyond-observed-support claims. This case establishes the existence, not the population prevalence, of failures in the measurement chain; it leaves acquisition versus redistribution unresolved.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.