acceptodds
Under review as a conference paper at ICLR 2027

Internal Attention: Growing Token Representations Without Growing the FFN or the KV Cache

Abstract

Increasing the width of a token's representation is expensive: parameters grow quadratically, dominated by the feed-forward network (FFN), and the KV cache grows linearly. We combine two ideas into an architecture that avoids both costs. First, inspired by Hyper-Connections, we split the token representation into a compute subspace and a memory subspace, where the FFN block is applied to the compute subspace only. Following Multi-Head Latent Attention, we also restrict the attention's key and value projections to the compute width. Quadratic costs in the number of parameters are therefore tied to the compute width and decoupled from the memory width. Second, we introduce a mechanism different from Hyper-Connections for moving information between the memory and compute subspaces: an attention internal to the token representation, which separates it into a fixed number of internal tokens and applies attention among them rather than across the sequence. An internal attention on the full vector followed by an FFN on the compute part forms a block that can be stacked, increasing depth without increasing the KV cache. Experiments on small autoregressive language models and integer multiplications show that the resulting architecture improves on both standard transformers and Manifold-Constrained Hyper-Connections (mHC) at equal KV cache, including in parameter efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.