Super Apriel: One Checkpoint, Many Speeds
Abstract
We introduce Super Apriel, a 15B-parameter supernet in which every decoder layer provides four trained mixer choices—Full Attention (FA), Sliding Window Attention (SWA), Kimi Delta Attention (KDA), and Gated DeltaNet (GDN). A placement selects one mixer per layer; placements can be switched between requests at serving time without reloading weights, enabling multiple speed presets from a single checkpoint. The shared checkpoint also enables speculative decoding without a separate draft model. The all-FA preset matches the Apriel 1.6 teacher performance; recommended hybrid presets span 2× to 10.7× decode throughput at 97% to 77% quality retention, with throughput advantages that compound at longer context lengths. With four mixer types across 48 layers, the configuration space is vast. A surrogate that predicts placement quality from the per-layer mixer assignment makes the speed–quality landscape tractable and identifies the best tradeoffs at each speed level. We investigate the convergence dynamics of the layer placement rankings. Rankings stabilize early at both scales, and while the ordering within the frontier placements stays somewhat volatile, the frontier set itself converges over distillation. Super Apriel is trained by stochastic distillation from a frozen Apriel 1.6 teacher, followed by supervised fine-tuning. We open-source the supernet weights, training code, vLLM serving code, and a placement optimization toolkit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.