Cross-Task Transfer via Multi-Role Reinforcement Learning
Abstract
Reinforcement learning with verifiable rewards is often used to train a language model for a single role that includes a problem-solving task and an objective based on answer correctness. Humans often learn differently; they decompose a problem or draw on relations that span examples, pursuing goals that are not directly tied to the correct answer to any single problem. Inspired by this, we ask whether training a language model on the same examples with auxiliary roles, alongside the primary one, can improve its test performance. We define each role by a task, what the model is shown and asked for, and an objective, how its output is rewarded. We formalize this setting as Role-Based Reinforcement Learning (RBRL), in which one model is trained in several roles. To evaluate our approach, we use GSM-Symbolic with a single distractor clause, which is known to substantially reduce LLM accuracy, and define two auxiliary roles: one that identifies the distractor clause within a problem and another that writes a method to solve every problem sharing a template. When the model is trained in these multiple roles, where the auxiliary roles include distractors but the primary role does not, the resulting model solves more problems with distractors at test time than when trained in the primary role alone. We interpret this gain in terms of cross-task transfer. Specifically, the training in the auxiliary tasks improves model performance on the primary task at test time. When distractors are also added to the primary role training, the model improves further, suggesting that the benefits of each role may be additive. We also observe the same additive pattern on reading comprehension tasks. An auxiliary role that states what a passage implies without actually asking any questions improves question answering on RACE-C and on five other datasets we did not train the model on.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.