acceptodds
Under review as a conference paper at ICLR 2027

CBSD: Cross-Batch Self-Distillation for Reinforcement Learning with Verifiable Rewards

Abstract

In language-model post-training with verifiable rewards, successful responses provide evidence for subsequent learning. GRPO-style response weighting, however, is restricted to current-batch gradients and may miss useful directions revealed by earlier successes. We introduce Cross-Batch Self-Distillation (CBSD), which separates rollout generation from evidence retention. An acting policy generates responses, while a non-acting carrier learns from verified successes across batches. The actor reads the carrier on its own sampled prefixes through on-policy distillation, obtaining token-wise guidance without replaying source responses or supplying them as context. A control-variate write and pre-write reading support stable training. Our analysis characterizes the ideal success-conditioned carrier and interprets its log-ratio read as token-level credit. In a binary-verifier family at a fixed actor, sufficient source sampling and reuse yield larger expected first-order gains than fresh-batch response weighting under matched total response and local KL budgets. With Qwen3-8B on SciKnowEval, CBSD achieves average maximum Avg@16, surpassing reported SDPO/GRPO by percentage points and leading in three of four domains. On 27 Physics prompts rarely solved initially, it exceeds GRPO by points at step 200.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.