MIDUS: Memory-Infused Depth Up-Scaling
Abstract
Depth Up-Scaling (DUS) expands pre-trained language models by inserting additional Transformer blocks, but duplicating FFN-heavy structures increases both parameter count and dense per-token computation. Sparse retrieval offers an alternative by decoupling stored capacity from dense computation, but efficient scaling requires avoiding large memory overhead while preserving the heterogeneous representations of individual attention heads. We propose Memory-Infused Depth Up-Scaling (MIDUS), which organizes memory around attention heads using distinct product-key spaces and shared value storage with head-specific projections. Under common retrieval weights, we show that shared value realization incurs a loss gap when different heads require different outputs, while MIDUS recovers the independent optimum under a shared factorization. Empirically, MIDUS remains competitive with strong DUS baselines while substantially reducing cost. On Llama-3.1-8B, MIDUS uses 8.3 fewer trainable parameters and 42% less peak training GPU memory than OpT-DeUS, together with higher generation throughput. Head-importance analysis further shows a more concentrated distribution across attention heads, consistent with the proposed head-wise memory design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.