Allocate Don't Operate: Online Data Allocation for Multi-Task LLM Post-Training
Abstract
Post-training a large language model rarely targets a single capability: a typical run mixes coding, tool-calling, mathematical reasoning, and instruction following. Classical multi-task learning studies this setting through the lenses of gradient interference and fairness, motivating methods such as gradient surgery and MGDA. These fit post-training poorly. Gradient-surgery methods must materialize per-task gradients at a cost that grows with both the number of tasks and the size of the model, while multi-objective methods target asymptotic balance in runs that stop well before convergence. We instead take the per-task quota of each mini-batch as the only control variable, casting multi-task post-training as online convex optimization: maximize cumulative policy improvement subject to a “no-task-left-behind” constraint. Assuming only directional concavity of the task-improvement surface, that constraint relaxes into a linear one computable from quantities already available at the current allocation. This yields Primal-Dual Curvature Matching (PDCM), a mirror-descent-like allocator driven by first-order estimates of how policy improvement responds to allocation, and unlike allocation methods developed for LLM pre-training it presumes no parametric model of training dynamics. By controlling batch composition alone, PDCM leads the uniform baseline on pass@1 by as much as 6.3% and consistently beats gradient-surgery and parametric mixing methods. Additionally, it scales to mixtures of up to 18 tasks, holding a margin a much as 5.3% over the uniform baseline with an increasing number of tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.