What Does RLVR Learn? Binding, Table Access, and Composition
Abstract
When reinforcement learning with verifiable rewards (RLVR) raises accuracy, what has the model learned? The same accuracy gain can reflect different outcomes: binding a new name to a familiar operation, fitting an arbitrary table of input–output pairs, or composing operations into unseen programs. We introduce Alien DSLs, executable languages whose symbol assignments and operation tables are sampled after pretraining, allowing us to evaluate these outcomes separately. Under outcome-only GRPO, exhaustive input evaluation supports partial binding between new names and familiar operations. Within the tested training budget, random-table accuracy plateaus near 30%, revealing a fitting bottleneck before transfer is assessed. Dense supervision largely resolves fitting, yet accuracy on the same observed cells drops from 96.5% to 13.7% when the query form changes. High training accuracy therefore does not ensure access to an observed table entry. For composition of renamed familiar operations, continued execution-trace supervision outperforms trace-initialized GRPO on held-out programs in smaller models. In the larger model, both recipes exceed an oracle per-program constant baseline on unseen operation pairs, unseen templates, and deeper programs. These transfer gains recur across two sequentially preregistered waves totaling sixteen fresh languages under a frozen training protocol. Together, these results show why fitting and transfer must be evaluated separately before interpreting an RLVR accuracy gain as evidence of a newly acquired capability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.