ReComp: Low-Rank Residual Compensation for Contextual Sparsity in Language Models
Abstract
Contextual sparsity, or dynamically computing a subset of neurons during inference, reduces multilayer perceptron (MLP) cost in Large Language Models (LLMs), but discards the output of the rest of the neurons, which is individually small but collectively substantial. We show that this output is predictable rather than lost, and evaluate whether part of a fixed compute budget is better spent on reconstructing it than computing more neurons. For a given token at layer , ReComp uses a low-rank neuron-importance predictor that selects the top- neurons to compute by importance. Then, a low-rank affine map estimates the combined output of the skipped neurons from the MLP input, fit via reduced-rank ridge regression (ridge-RRR) and distilled against the dense model with the objective having a KL-divergence term and an optional Jacobian-lens term. Across five models from 0.5B to 14B parameters, reconstruction beats spending the same amount of computation on more neurons at every budget we test with our selector. For our predictor, the compensator helps even before distillation. After distillation, it also improves other selectors; for example, when added to TEAL, it lowers relative perplexity on Qwen2.5-7B from 1.42 to 1.16 at the same arithmetic. At 25% of compressed-layer MLP arithmetic, ReComp lowers relative WikiText-2 perplexity on Qwen2.5-7B from 1.44 for the best of TEAL, WINA, and R-Sparse to 1.16, and beats all three on every 7-8B model tested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.