acceptodds
Under review as a conference paper at ICLR 2027

LatticeLoop: Learning Loop-Specific Weights from Shared Low-Bit Codes

Abstract

Looped Transformers apply one stack of core layers several times. Full weight tying keeps the model small, but every loop must apply the same weights. Independent weights for each loop remove this restriction, but the core storage grows with the number of loops. Partial untying and per-loop LoRA add less storage, but they limit the change between loops to a few rows or to a low-rank update. Block-scaled low-bit formats such as NVFP4 already split every weight tensor into two parts: 4-bit codes set the pattern of each block of 16 weights, and 8-bit block scales set its magnitude. LatticeLoop shares the codes across loops and learns one table of block scales for each loop. Each additional loop costs 0.5 bits per core weight, one ninth of an independent copy, and every loop runs as one standard block-scaled matrix multiplication. In language model pretraining at matched storage, LatticeLoop closes 92.5% of the loss gap between full tying and independent weights, while partial untying and per-loop LoRA close less than a third of the gap. The advantage holds up to 12 loops and 80B training tokens, and decoding speed stays close to full tying.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.