Reward-Proportional Routing of LoRA Experts for Continual Learning in Large Language Models
Abstract
Continual learning (CL) of large language models (LLMs) commonly assigns task-specific low-rank adaptation (LoRA) experts to a frozen backbone and trains a router that selects experts for each input without task identity. Although the number of experts can vary for different inputs, it is inappropriately fixed in advance, as existing routers are only trained to identify the best experts. A small change in the loss can swap the entire selection, which degrades performance on previous tasks with their experts frozen. This paper proposes a method of reward-proportional routing (RPR) that trains the router toward a distribution over expert combinations that is proportional to their reward. The distribution changes continuously with the loss, and reading it up to a cumulative probability mass determines both the number and the composition of experts for each input. RPR learns the distribution with a generative flow network (GFlowNet) and freezes one expert per task. On three CL benchmarks with a frozen T5-large, RPR achieves the highest mean accuracy on all eight task orders, with against for the second-best method and the lowest average forgetting on the 15-task Long Sequence benchmark. On tasks never trained on, RPR reaches against for the unadapted backbone, whereas the compared methods produce labels of previous tasks and reach at most .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.