acceptodds
Under review as a conference paper at ICLR 2027

TwIST: Rigging the Lottery in Transformers via Zero-Shot Subnetworks

Abstract

We introduce , a novel distributed system for efficient training and deployment of Transformer-based models. unifies two important paradigms: distributed training and pruning, to induce an inference time property that allows us to randomly sample competitive subnetworks on the fly. trains randomly generated subnetworks in parallel, periodically aggregating and resampling, yielding high-performance subnets that require no fine-tuning or search heuristics. This enables robust, zero-cost pruning at deployment, achieving perplexity scores close to state-of-the-art post-training methods while bypassing their post-training overhead (e.g., calibration, Hessian inversion). Crucially, even when post-training baselines operate on a stronger dense checkpoint than , they still lose at aggressive sparsity, often by orders of magnitude, evidence that designing training around subnet structure produces better sparse models than searching for them after the fact. We validate this across three architectures (GPT-2 and GPT-2 XL, with multi-head attention and Qwen3-0.6B with grouped-query attention) over a wide range of text-generation tasks. 's advantage emerges under aggressive pruning (e.g., 50%+ sparsity), where it significantly outperforms baselines. For example, on GPT-2/WikiText-103 it achieves 23.14 PPL while the closest baseline follows at 31.64. At the most aggressive Qwen3-0.6B pruning setting, reduces perplexity by approximately \(2.5\)–\(4.6\times\) relative to the strongest evaluated baseline, as those baselines collapse entirely. Furthermore, as a structured pruning method, produces smaller, dense matrices, translating to tangible inference speedups and memory savings on inference server deployments (e.g., CPUs) that lack sparse computation support. We provide the complete implementation https://anonymous.4open.science/r/twist2-21E6here.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.