acceptodds
Under review as a conference paper at ICLR 2027

Nesting Is All You Need: Recursive Memory Plane with Latent Specialization

Abstract

Most language models treat depth as disposable computation: intermediate representations are transformed and discarded until only a terminal state remains. We introduce **Recursive Memory Plane**, a language-model architecture built around nested specialization: Primary (P), Intermediate (I), and Association (A) cells are structurally specialized for local, contextual, and global computation, forming a persistent P→I→A hierarchy whose entire representational trajectory remains jointly addressable. Rather than recur only over the terminal hidden state, each recursive step performs a **reason–readdress–integrate cycle**: a reasoning state (R) queries the full memory plane, then a discourse state (D), conditioned on the updated reasoning state, independently re-queries the same plane to form the next prediction state. **LatentBank** further adds compact, persistent I- and A-specialist states, preserving nested specialization alongside recursion. This design yields an unexpected insight: **recursion is most valuable during training, while shallow inference performs best**. At 220M scale, LatentBank adds only 0.23% parameters; training with three recursive steps and inferring with one achieves a **34.66% raw mean across nine benchmarks**, versus **33.87% for the strongest parameter matched baseline**. Eliminating inference recursion still reaches 34.52%, whereas training without recursion falls to 32.99%. These results suggest a different design principle: **nested specialized computation, preserve depth as memory, recursively re-address the full representational trajectory, and use deep training to produce powerful shallow inference.**

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.