TD-Grokking: Learning from Zero-Reward Problems by Training-Time Decomposition
Abstract
Reinforcement learning with verifiable rewards (RLVR) has advanced LLM reasoning, but it cannot learn directly from difficult problems for which every sampled solution fails: their uniformly zero rewards provide no contrastive training signal. We propose TD-Grokking, a training-time decomposition framework for learning from these zero-reward problems. Guided by a reference solution, it rewrites the local reasoning requirements of each original problem as self-contained subproblems with verifiable answers. Joint training on the original problems and their subproblems allows the model to learn from verifiable local signals while continuing to practice complete solutions. Across all 15 model–benchmark combinations spanning three model families and five mathematics and science reasoning benchmarks, TD-Grokking outperforms vanilla GRPO by 4.30% on average. It also consistently improves performance across four medical question-answering benchmarks. Further analysis shows that training on subproblems alone enables 77 of 507 previously zero-reward original problems to produce at least one verified correct solution, despite never training directly on those problems. These results support training-time decomposition as a practical way to obtain useful outcome-reward signals from difficult reasoning tasks. Our code and datasets are available at https://anonymous.4open.science/r/TD-Grokking-6567/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.