acceptodds
Under review as a conference paper at ICLR 2027

SGA-Sparse: Training Native SwiGLU Gates for Activation-Sparse and Low-Bit LLMs

Abstract

Activation sparsity can reduce FFN weight traffic during single-user LLM decoding, but active channels must be identified early enough to avoid fetching unused weights. In pretrained SwiGLU models, training-free routing with the native gate performs substantially worse than routing with the up activation at high sparsity. We show that this ordering reverses after sparsity-aware continued pretraining. To the best of our knowledge, we present the first systematic study that converts pretrained SwiGLU LLMs into fixed-budget activation-sparse models by training the existing gate projection as the router. Our method, SGA-Sparse, updates all model parameters through continued pretraining while preserving the original SwiGLU activation and introducing no auxiliary router parameters. After approximately 30B training tokens, LLaMA models from 1B to 8B achieve uniform per-token 75–85% FFN channel sparsity while remaining within 0.2–2.7 pp of their original dense checkpoints in average zero-shot accuracy. We further study two complementary QAT recipes: continued-QAT largely recovers the performance at 4-bit weight quantization, while joint-QAT performs better for the more aggressive sub-4-bit target. To compare activation sparsity with dense low-bit quantization, we translate the sparsity-induced reduction in weight traffic into an equivalent bit-width. Under this common idealized metric, our 2.5-bit-equivalent models achieve average zero-shot accuracy comparable to dense 3-bit QAT models, while our 1.5-bit-equivalent models approach dense 2-bit models and outperform dense ternary models by 1.3 pp on average. These results establish native-gate training as a practical approach for combining high FFN sparsity with low-bit weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.