Lossy Projection, Exact Expansion: Retrofitting Pretrained Gated Feed-Forward Blocks
Abstract
Weight movement can limit autoregressive language-model inference, motivating gated feed-forward architectures that reuse weights across computational paths. Masked Gated Linear Units (MGLUs) construct gate and value projections using complementary masks over a shared matrix, with multiple masks providing additional routes without duplicating that matrix. Converting pretrained blocks into this representation can incur branch-reconstruction error, while initializing additional routes can further alter the projected function. We propose Projection-Aware Route Expansion (PARE), which separates dense-to-shared projection from function-preserving multi-mask expansion. Our analysis characterizes the reconstruction limits of complementary sharing and establishes an exact embedding of the projected single-mask parent through repeated masks and a rescaled down projection. Mask logits are initialized with matching signs but distinct magnitudes across routes, yielding identical binary masks at entry while permitting differentiation during recovery. In the controlled 159M experiments, four-mask PARE reaches 15.65 perplexity versus 16.06 for fixed single-mask recovery at the same recovery-token budget. Within the same four-mask architecture, PARE also achieves lower endpoint perplexity than quota-preserving and reconstruction-calibrated diverse initializations. These endpoint advantages persist when the frozen checkpoints are evaluated on source-excluded FineWeb-Edu and C4. Further experiments on TinyLlama-1.1B and Gemma-3-1B, PARE also achieves lower endpoint perplexity than fixed single-mask recovery. A matched four-mask comparison on TinyLlama additionally favors PARE over diverse initialization. These results support function-preserving route entry for improving conversion fidelity at fixed token budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.