acceptodds
Under review as a conference paper at ICLR 2027

Gen-rl: Rewarding Reasoning That Transfers

Abstract

In standard reinforcement learning with verifiable rewards (RLVR), a reasoning chain is rewarded for solving the particular problem that elicited it. But the reasoning we ultimately want models to learn should transfer: a broadly reusable strategy and an idiosyncratic solution should not be equally valuable. We introduce generalization-rl (*gen-rl*), which uses in-context learning as a direct signal for transfer. The model first produces a response to one problem, which is left in context as a demonstration for a set of subsequent problems. The first response is rewarded according to how well it enables downstream problem solving via in-context learning, while the downstream responses are themselves trained to make effective use of the demonstration and produce correct answers. We derive a principled policy-gradient estimator for learning this behavior end-to-end. We study transfer in one concrete form, easy-to-hard generalization: whether reasoning learned on problems of one difficulty extends to harder ones. Across multi-digit multiplication, knights & knaves, and standard competition math benchmarks, gen-rl outperforms standard RLVR both on problems as difficult as the training problems and on harder ones from held-out difficulty tiers. Gen-rl's improvement over standard RLVR is largest on the harder tiers, suggesting that it promotes reasoning strategies that extrapolate beyond the training distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.