MoEcraft: Trajectory-Based Sparse Upcycling of Instruction-Tuned Small Language Models
Abstract
We study sparse upcycling of instruction-tuned small language models into mixture-of-experts (MoE) models without returning to large-scale continual pre-training. MoEcraft reuses checkpoints from a single on-policy reinforcement-learning (RL) trajectory to initialize experts, followed by lightweight joint router-expert calibration. It combines all-linear low-rank adaptation, feed-forward adapters from eight checkpoints, and top-2 routing with calibration on 512 examples. In the canonical Qwen3-1.7B run, the chain scores 52.35 at the dense RL endpoint, 51.84 after sparse assembly, and 56.95 after calibration on Overall13, the mean over 13 benchmarks. Across three RL trajectories, calibration adds 3.60-5.11 points, and final performance averages 56.14 +/- 0.72 (mean +/- sample standard deviation). A single-run supervised fine-tuning trajectory control using the same sparse construction and calibration gains 2.50 points from assembly but loses 0.64 during calibration, driven by a summarization regression. The stage-wise pattern also recurs with a second RL optimizer and on Llama and Falcon models. The tested pipelines complete within one day per model on a single 96GB GPU. Under a fixed serving setup, the 1.7B-derived MoE achieves 1.54x/1.22x the output throughput of released dense Qwen3-4B at batch sizes 1/4 with 2.76B active parameters, but has 9.10B total parameters and a larger resident footprint (30 versus 20 GiB). These results characterize the utility and performance-footprint trade-offs of reusing RL trajectories for post-alignment sparse upcycling in the tested settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.