SGRPO: Optimizing State-Space Language Models with State-Isolated Group-Relative Policy
Abstract
article iclr2027_conference,times math_commands.tex amsmath,amssymb graphicx booktabs hyperref url SGRPO: Optimizing State-Space Language Models with State-Isolated Group-Relative Policy Anonymous Author(s) Affiliation withheld for double-blind review %\iclrfinalcopy % document \maketitle abstract Group-relative policy optimization (\GRPO) assigns each sampled response an advantage by standardizing its reward against a group of rollouts generated from the same prompt. The derivation of this estimator requires the rollouts in a group to be independent samples from the current policy—a condition Transformer policies satisfy, since each rollout re-encodes the prompt and retains no state between generations. State-space policies such as Mamba break this condition: because they carry state recurrently, decoding a group's rollouts sequentially against one shared generation cache, an efficiency choice that avoids re-encoding the prompt, lets each layer's recurrent state and causal-convolution buffer carry forward from one rollout to the next. Each rollout after the first therefore begins not from the clean post-prompt state produced by the prompt alone, but from the state left behind by the preceding rollout. Under a non-degenerate reward map, this carryover gives the estimator an order-dependent bias that, through the shared group baseline, reaches even the uncontaminated first rollout. The correction, state isolation, restores each layer's recurrent state and convolution buffer to their post-prompt values before every rollout, reinstating exact independence at zero additional FLOPs and leaving the group-relative objective unchanged. packages this correction with token-level normalization and future-KL token weighting. The correction is worth adopting for its exactness and negligible cost alone, regardless of its effect on downstream accuracy. abstract %references %iclr2027_conference document
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.