acceptodds
Under review as a conference paper at ICLR 2027

RANDOMLLM: HOW LITTLE RANDOMNESS DOES A STRONG LOTTERY TICKET NEED?

Abstract

Strong-ticket training freezes a network's weights at random values and learns only a mask over them. The standard construction for language models samples an independent random weight for every projection, so the random source holds as many independent scalars as the model has connections. We propose RandomLLM, which instead slices every random weight from a single rank- random pool and keeps a separate multivalued mask for every projection. We find that the randomness of the random weight has little influence on the strong ticket that mask learning produces: at source rank , two frozen vectors supply independent scalars, about times fewer than one per connection, and the resulting model reaches perplexity against and against mean zero-shot accuracy at 3B, while a controlled three-seed experiment finds no resolvable effect from either sharing the pool or constraining its rank. Moreover, instead of accelerating this paradigm with dedicated hardware; we give a general-purpose GPU implementation that avoids rebuilding the mask at every forward pass, reducing prefill and decode latency by and and training time by .

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.