Beyond Correctness: Learning Robust Reasoning via Transfer
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has recently strengthened LLM reasoning, but its focus on final-answer correctness leaves a critical gap: equally correct solutions can differ in the usefulness of their intermediate reasoning. We adopt a simple philosophical view—good reasoning should remain useful beyond the mind that produced it—and treat reasoning as a form of meaning transfer that can be assessed through truncation and continuation by another model. Building on this principle, we introduce Reinforcement Learning with Transferable Reward (RLTR), which rewards reasoning prefixes that support correct continuation by a separate model. This objective turns the usefulness of intermediate reasoning to another model into an auxiliary training signal for improving the generator itself. Our approach improves majority-vote accuracy while improving final-answer accuracy, and it reaches comparable performance in substantially fewer training steps. For example, RLTR gains a +5.8%p in Maj@64 on AMC23 compared to RLVR, and RLTR matches RLVR’s average accuracy with roughly 2.5× fewer training steps on MATH-500. Extensive analyses further support RLTR’s effectiveness, showing improved prefix utility for held-out receivers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.