acceptodds
Under review as a conference paper at ICLR 2027

Latent Lattice Language Models

Abstract

Standard autoregressive transformers route all per-token computation through a single active latent. This latent is responsible for querying and processing the "memory" from the past steps, forming new memory (updating the KV cache) for the future steps, producing residual updates to its own state, and finally decoding into the next token. The memory is a write-once, read-only log of the latent's past states, while all active computation still rests on that one latent vector. Recent work increases the computation performed by this single latent either along depth - by adding more layers or looping - or along time, by letting it persist across steps rather than being decoded at each one. However, the model remains bottlenecked, relying on a single latent to perform every role. We instead ask: how can we increase the computational latent capacity available to a language model, and what does that buy us? To that end, we introduce an additional horizontal latent at each Lattice layer, updated continuously along the temporal dimension. Together with the standard vertical latent that runs along the network's depth axis, these horizontal recurrent states form our Lattice of Latents. Unlike hybrid Transformer-SSM architectures, where recurrent state typically remains internal to a sequence-mixing operator (primarily focused on replacing softmax-attention and KV cache), our horizontal latent acts as an explicit companion computational state to the vertical latent, jointly driving attention, KV-cache writes, MLP computation, and decoding. We propose several mechanisms to increase the expressivity of these horizontal latents while retaining token-parallel teacher-forced computation within each pass. Across pre-training budgets of 10B-30B tokens, Lattice models consistently improve language-modeling quality and downstream transfer. After pre-training on 20B FineWeb-Edu tokens, our 406M-parameter Latent Lattice Language Model lowers perplexity from to relative to a parameter and compute-matched Transformer, while also matching a 609M-parameter, 40-layer Transformer that uses as many parameters, the KV-cache footprint, and as many inference FLOPs. The gains also extend to downstream evaluation and mathematical reasoning after finetuning. Project Page: https://latent-lattice-language-model.netlify.app/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.