ReFiT: Recovering Block-Sparse FFN Quality at No Extra Inference Cost
Abstract
Long-context tasks require processing many input tokens before generation, making efficient LLM prefill increasingly important. Sharing FFN neuron sets across tokens preserves efficient batched computation, but limits adaptation to each token and degrades quality. Individual-neuron energy scores omit interactions between discarded contributions, while a block-constant correction cannot recover token-dependent errors. We propose ReFiT, a closed-form refit of the FFN down-projection weights to reconstruct missing output from surviving activations under a fixed selection policy. The update performs ridge regression of the discarded output on retained activations, regularized toward the original weights. These activations remain token-specific under a shared mask, allowing one fitted down-projection matrix to produce different corrections for different tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.