Task Vectors in the Loop: Online Geometric Coordination for Multi-Task RL for LLM
Abstract
Reinforcement-learning (RL) post-training has advanced the reasoning and problem-solving capabilities of large language models, but extending these gains across heterogeneous tasks requires balancing cross-task transfer with task specialization. Mixed RL and post-hoc model merging approach this challenge from opposite directions: joint training allows transfer but risks interference, whereas independent training preserves specialization but leaves task-specific updates to be reconciled afterward. Rather than treating specialization and joint learning as competing choices, we use brief specialization to reveal task structure that guides joint optimization. We introduce Task-Vector-Guided Reinforcement Learning (TVG-RL), which turns short task-specific warm-ups into a geometric prior for multi-task RL. Starting from a common initialization, TVG-RL extracts task vectors and decomposes their structure into shared and task-private subspaces. It then trains a single policy from the original base, using this decomposition to project and rescale subsequent RL gradients. Shared and private components receive distinct treatment, while a complementary component permits updates outside these task-conditioned shared–private subspaces. The resulting method requires neither fully trained experts nor a final model-merging stage. Experiments across eight heterogeneous benchmarks show improvements over direct mixed-data RL on most tasks, alongside gains over the evaluated post-hoc model-merging baselines. These findings support a role for task specialization beyond producing standalone experts: the geometry it reveals can help a shared policy learn multiple capabilities together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.