The Lottery Circuit Hypothesis: Sparse Mechanisms from Randomness
Abstract
Interpretable AI, and mechanistic interpretability in particular, aims to uncover the internal computations of language models, with circuit discovery remaining an important approach. We propose the Lottery Circuit Hypothesis: sparse, task-functional circuits can be constructed by zero-based masking of random-weight transformers, rather than merely extracted from computations established through pretraining. This reveals a possibility in circuit discovery that bears on what we call the provenance assumption: that a sparse subgraph reproducing a target task behaviour thereby localises computation already encoded in the pretrained weights. Motivated by the strong lottery ticket hypothesis, we perform circuit discovery beyond pretrained models and find that DiscoGP can construct sparse, high-performing weight-and-edge circuits in randomised transformers. We further demonstrate concealment: task-agnostic, small-magnitude fills obscure the trace of zero-based weight pruning while preserving task performance. This extends the perturbation perspective of neural thickets from neighbourhoods of pretrained models to neighbourhoods of masking-constructed subnetworks within random models. We give a theoretical account of what we call margin synthesis by linking this constructive capacity to mask granularity: in a simplified additive model, each edge contributes a sum of independent, zero-mean Gaussian terms, normalised by , to the task margin. Weight pruning can retain favourable terms within an aggregate and attain , whereas edge pruning selects among aggregates of magnitude . Together, these results call for a rethinking of circuit discovery: a discovered circuit may be a mask-induced construction, and interpreting it as a map of pretrained computation requires additional treatment and evidence.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.