Capability Budgeting: Acquiring Mathematical Reasoning While Retaining Capabilities in Multilingual Continual Pretraining
Abstract
Multilingual continual pretraining should acquire target skills while retaining capabilities that remain useful. We study this objective under a fixed token budget through Capability Budgeting: allocating data by role and assessing mathematical acquisition against explicit per-capability loss budgets. A controlled study jointly adapts Thai, Bengali, and Swahili on Qwen3-4B-Base and Qwen3-8B-Base, comparing three allocations with three training seeds each and 16.384M tokens per run. All 18 trained models improve target-language average mathematics accuracy by 2.27–4.27 percentage points, full-test English GSM8K, and Swahili Global-MMLU. Mean GSM8K gains range from 2.70 to 3.92 points at 4B and 1.47 to 1.84 at 8B; all six code-evaluated 4B models gain 6.10–7.93 points on full-test HumanEval. Every matched-seed comparison at both sizes contains a trained candidate combining mathematical gains with all nine measured-score retention requirements. At a fixed combined English-pool share, exchanging mathematical for general text changes the distribution of target-language and broader task outcomes. These differences matter for selection: retention requirements change the highest-mathematics choice in every 4B seed comparison. The results demonstrate mathematical acquisition with budgeted capability retention, and show how data roles and individual task outcomes inform the choice of an adaptation recipe.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.