Lattice: Learning to Efficiently Compress the Memory
Abstract
Attention mechanisms have revolutionized sequence learning but remain bottlenecked by quadratic computational complexity. This paper introduces Lattice, a novel recurrent neural network (RNN) mechanism that leverages the inherent low-rank structure of key-value matrices to efficiently compress contextual information into a fixed number of memory slots, achieving sub-quadratic complexity. By framing this compression as an online optimization problem, we derive a dynamic memory update rule based on a single gradient descent step. The resulting recurrence features a state- and input-dependent gating mechanism, offering an interpretable and theoretically grounded memory update process. The core innovation is the orthogonal update: each memory slot is updated exclusively with information orthogonal to its current state, hence incorporating only novel, non-redundant data to minimize interference with previously stored information. To ensure hardware-efficient training, we derive specialized chunk-wise parallelization forms tailored to this non-linear recurrence. Empirically, Lattice outperforms strong baselines on language modeling and associative recall tasks across diverse context lengths and model sizes, achieving superior memory efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.