Where, What, and When: Structured Intra-Response Advantage Redistribution for LLM Reasoning
Abstract
Reinforcement learning with verifiable rewards has substantially improved large-language-model reasoning, yet Group Relative Policy Optimization (GRPO) typically applies the same response-level advantage to all tokens in a sampled response. We introduce Where–What–When (WWW), a structured intra-response advantage redistribution framework that retains GRPO's group-relative advantage estimator while differentiating credit across reasoning segments. Where supplies a structural position prior, What uses segment-level surprisal evidence to select candidate segments, and When uses their disagreement to adapt the intervention strength and selectively invoke Beam–Monte Carlo Verification (BMCV) for outcome-sensitive calibration. Response-wise mean normalization controls the average weight scale before the redistributed advantage is used in the clipped policy objective. On Llama-3.2-3B-Instruct, WWW improves the best observed accuracy over our matched GRPO baseline by 2.20–3.34 percentage points across GSM8K, MATH500, AIME24, and AIME25. On Qwen2.5-7B-Instruct, within the first 500 training steps, WWW improves the best observed accuracy over GRPO by 2.84–6.70 points across the same benchmarks. The lightweight WWW variant without BMCV adds only 1.99% end-to-end step-time overhead, while Full WWW adds 12.69%. We will release our code and model weights.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.