From Target Shift to Covariate Shift: How RL Post-Training Enables Transfer to New Tasks
Abstract
Reinforcement learning (RL) post-training has substantially improved the reasoning capabilities of autoregressive language models, but whether it can elicit solutions to tasks not explicitly targeted during pre-training remains unclear. We study this question through the canonical problem of learning parities over uniform input bits. Pre-training data consist primarily of short answers to a parity task, together with rare chain-of-thought (CoT) demonstrations that include intermediate computations. During RL post-training, the reward instead targets a distinct parity function that is uncorrelated with the pre-training target. Experimentally, we observe that RL post-training can achieve nontrivial accuracy on the new task, escaping apparent computational barriers, while supervised fine-tuning under the same information is ineffective. During the course of RL training, networks undergo sudden phase changes in both accuracy and generation length, and can converge to suboptimal strategies. By developing an effective theory of the RL dynamics, we characterize some of these suboptimal strategies, and link them to the formation of low-dimensional features that can be learned efficiently, yet are predictive of the new target over only some parts of the input space. We explain how the RL loss encourages the discovery of these features by turning the target shift into a covariate shift, allowing the model to repurpose pretrained CoTs on inputs where the two targets coincide.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.