CrossLayerTopK: Dynamically Allocating Sparsity Across Layers in Transcoders
Abstract
Cross-Layer Transcoders (CLTs) decompose the activations of large language models into human-interpretable features and trace how these features interact across layers to form circuits, making them a key tool for mechanistic interpretability. However, training CLTs effectively and efficiently is hard. Beyond the usual trade-off between sparsity and reconstruction quality, a CLT has to model all layers at once, which adds a further challenge. The common JumpReLU recipe controls sparsity only indirectly through a penalty, while the direct extension of TopK selects a fixed number of latents in every layer. Yet reconstruction difficulty differs widely across layers, and intermediate layers are much harder than others, so giving every layer the same budget is suboptimal. We propose \CLTK, which sets a single budget for the whole model and dynamically allocates it across layers during training. In a comparison with JumpReLU and other TopK variants on Qwen3-8B-Base, this gives lower FVU while using fewer active latents than JumpReLU. The better reconstruction carries over to circuit tracing, yielding more complete circuits with higher replacement and completeness scores. The fixed per-token budget of \CLTK also makes it easy to pair with sparse decoding, which cuts both computation and communication and gives a speedup per training step. We further show that \CLTK follows scaling laws in both latent size and sparsity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.