Task Specialization Fine-Tuning for Contextual Reinforcement Learning
Abstract
Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initial performance, followed by fine-tuning multiple policies for task specialization. This new paradigm, however, introduces unique challenges, such as heterogeneous marginal returns and sample inefficiency. This raises a critical research question: given a pretrained policy and a constrained budget, *how much* fine-tuning should each task region receive to enable sample-efficient CRL? To this end, we propose *Task Specialization Fine-Tuning (TSFT)*, an online framework that formalizes budget allocation as a Markov decision process and uses a simple parametric function to predict fine-tuning performance. TSFT leverages dynamic programming to derive the optimal allocation, with a receding-horizon variant that trades optimality for efficiency. Extensive experiments across diverse decision domains, including combinatorial optimization, continuous control, and LLM fine-tuning, demonstrate that TSFT significantly outperforms baselines in task coverage and approaches oracle performance. Our work charts a new direction for model-based CRL, aligning with the modern pretrain-finetune era.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.