AtomShare: One Shared Atom Set Across a Mixture-of-Experts Stack for Budget-Constrained Weight Compression
Abstract
A sparse mixture-of-experts stack stores one weight matrix per expert at every depth, and its stored size therefore grows with the expert count while only a few experts run per token. However, existing compressors fit each matrix in a basis of its own and price its error by how far that matrix moves; they therefore cannot reach the redundancy that recurs across the experts of a layer and they spend the budget against a distance the router never applies. We propose AtomShare, which stores a whole stack as one atom set shared across depths, carried into each layer by a norm-preserving map, with every expert held as a sparse code against that dictionary. AtomShare achieves this by fitting one atom set against every expert of the stack in each layer's recorded input geometry, by setting each unit's code length from the deviation it delivers at the layer output on the tokens the router sends it, and by dividing one stored-size budget across the layers by the deviation each layer's injection propagates to the end of the stack. Sharing the dictionary is what makes code length a single divisible pool, and the atoms, maps and codes are therefore fit together as one alternation rather than matrix by matrix. On an eight-layer, forty-eight-expert weight stack AtomShare reaches a retained score of 0.5483 against 0.5150 for the strongest per-expert baseline while holding fewer bits per parameter, and its compression ratio rises with the expert count, from 6.149 at eight experts to 8.385 at sixteen, where every per-expert store but one holds its own ratio fixed across that range. A mechanism ablation read at the routed output puts the recorded input geometry first: removing it raises the delivered deviation by 67 to 93 per cent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.