When RL Suppresses Its Own Vocabulary: Recovering Reasoning Diversity in Puzzle-to-Math Transfer
Abstract
Reinforcement learning using verifiable rewards (RLVR) improves LLM reasoning, but the conditions under which it transfers across domains, and why it does so, remain under-explored. We study cross-domain transfer in two model families whose SFT and RL post-training stages use only constraint-satisfaction puzzles, with no mathematics problems in the post-training data, using a reasoning primitive-level analysis framework: a 9-class span classifier that segments chain-of-thought reasoning into primitives and tracks how they are sequenced across recipe stages and domains. We find that puzzle SFT induces a reasoning-primitive vocabulary, and that vanilla GSPO then commits the model to longer uninterrupted chains of reasoning, yielding pass@32 gains on OlymMATH-Hard in both families. But the same training also suppresses backtracking across families and in both domains. To address this, we propose a novelty bonus that rewards diverse correct rollouts using perplexity under the reference model as a signal, designed to preserve the backtracking present in the SFT checkpoint. It brings backtracking back toward its SFT level, and in a controlled A/B test against vanilla GSPO (matched data, algorithm, training budget) it adds a further pass@32 gain on hard mathematics. The end-to-end recipe raises hard-math pass@32 by 20 points on OlymMATH-Hard for OLMo3-7B and 7 points on HMMT for Qwen3-4B without adding any mathematics problems in the post-training stages studied here.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.