acceptodds
Under review as a conference paper at ICLR 2027

Tandem Reinforcement Learning with Verifiable Rewards

Abstract

Reinforcement learning with verifiable rewards (RLVR) has driven substantial gains in language-model reasoning. However, models fine-tuned with RLVR need not remain compatible with weaker partners: gains achieved alone may erode under collaboration, while policy drift may place reasoning beyond a weaker model's predictive reach. To address both problems, we introduce Tandem RLVR (TRLVR), bringing the recently proposed tandem-training paradigm beyond proof-of-concept settings and into modern RLVR. In TRLVR, a senior learns from rewarded trajectories co-generated with a frozen junior, while the standard GRPO objective is applied only to senior-emitted tokens. Training Qwen3-4B-Instruct on competition math, TRLVR performs on par with a matched GRPO senior when reasoning alone. We characterize compatibility through the complementary lenses of handoff robustness and junior legibility. Under reasoning-step handoffs, the TRLVR senior retains nearly all of its solo performance with the junior, reducing the communication tax to a negligible loss. The senior's reasoning is also more legible to the junior: its token distribution stays more closely anchored to the junior's, and the junior predicts the senior's chain-of-thought more readily token by token. An ablation that explicitly regularizes GRPO toward the junior does not recover the same combination of capability and compatibility. These results provide initial evidence that tandem rollout structure is a promising direction for keeping RLVR reasoning gains accessible to weaker partners.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.