acceptodds
Under review as a conference paper at ICLR 2027

MIDUS: Memory-Infused Depth Up-Scaling

Abstract

Depth Up-Scaling (DUS) expands pre-trained language models by inserting additional Transformer blocks, but duplicating FFN-heavy structures increases both parameter count and dense per-token computation. Sparse retrieval offers an alternative by decoupling stored capacity from dense computation, but efficient scaling requires avoiding large memory overhead while preserving the heterogeneous representations of individual attention heads. We propose Memory-Infused Depth Up-Scaling (MIDUS), which organizes memory around attention heads using distinct product-key spaces and shared value storage with head-specific projections. Under common retrieval weights, we show that shared value realization incurs a loss gap when different heads require different outputs, while MIDUS recovers the independent optimum under a shared factorization. Empirically, MIDUS remains competitive with strong DUS baselines while substantially reducing cost. On Llama-3.1-8B, MIDUS uses 8.3 fewer trainable parameters and 42% less peak training GPU memory than OpT-DeUS, together with higher generation throughput. Head-importance analysis further shows a more concentrated distribution across attention heads, consistent with the proposed head-wise memory design.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.