Growing and Transferring Pathways for LLM Reinforcement Learning
Abstract
Reinforcement learning (RL) can substantially improve the reasoning capabilities of large language models, but an effective training run requires considerable resources. Reusing the computation acquired through RL across models would extend the value of that training. We propose GrowRL, an RL method that organizes learned updates into compact computational pathways through structural growth. GrowRL freezes the pretrained backbone and adds an independent causal attention branch and a feed-forward network branch at each Transformer layer. Zero-initialized output projections preserve the pretrained policy at initialization, and RL updates only the added pathways. The learned pathways remain explicit after training, allowing transfer to structurally compatible target models and subsequent adaptation with the target backbone frozen. The architecture also supports sequential transfer of additional pathways. By keeping the learned attention and feed-forward computation separate from the pretrained weights, GrowRL provides an explicit structure for preserving and reusing the outcome of an RL run. Experiments across five backbones show competitive mathematical reasoning and out-of-distribution performance with a small number of trainable parameters. Learned pathways provide immediate gains across all three transfer directions, with performance after only 10 target-side updates already close to that after 100 updates in every case. These results support computational pathways as parameter-efficient and reusable units of reinforcement learning adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.