Task Reformulation Enables Reinforcement Learning From Hard Reasoning Problems
Abstract
Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but it cannot learn from problems that are too difficult to solve under the current policy, since these yield no reward signal and therefore no gradient. We propose a simple solution based on task reformulation. We transform challenging open-ended problems into cognitively simpler variants (such as multiple-choice and cloze formats) that preserve the original answer while reducing the effective search space and providing denser learning signals. These variants span a spectrum from discriminative to generative tasks, so a model can first learn from the structured, easier formats and transfer that knowledge back to the original problems. Building on this insight, we introduce Cog-DRIFT, a framework that automatically constructs reformulated variants and organizes them into an adaptive curriculum, progressing from easier to harder formats as the model improves. This lets the model learn from problems that previously yielded zero signal under standard RL post-training, without distilling reasoning from any stronger external model. Cog-DRIFT improves on the originally unsolvable hard problems (+10.11% for Qwen and +8.64% for Llama in absolute gains) and also generalizes to held-out datasets. Across two models and six challenging reasoning benchmarks, it outperforms standard GRPO and strong guided-exploration baselines that supply stepwise cues or partial solutions on average, improving the average accuracy by +4.72% (Qwen) and +3.23% (Llama) over the strongest baseline in our main comparison while needing only the gold answer. We further show that Cog-DRIFT improves pass@k at test time and that the curriculum improves sample efficiency. Cog-DRIFT outperforms GRPO when GRPO is given 8 times more rollouts or 4 times longer generations. Overall, our results highlight task reformulation and curriculum learning as an effective paradigm for overcoming the exploration barrier in LLM post-training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.