PixelMem: Persistent Images for conteXt Efficient Language-model Memory
Abstract
For language models to support long-term user interactions, they need to retain and reuse user information, past events, and procedural knowledge. These memories require storage space, and reading them back into the model occupies its limited input capacity. Rendering text as images can shorten model inputs, but may enlarge the storage cost. This motivates a memory format that supports both compact storage and short model inputs. We introduce PixelMem, which encodes information from text memories into the pixel values of images. A trained visual reader maps each image to a small number of continuous memory slots, which the language model uses together with the question to generate an answer. The image size remains fixed, while the reader determines how many input positions the memory occupies. We evaluate storage, input length, and answer accuracy on four types of tasks, ranging from simple facts to detailed procedures. Each question is answered from a single memory. With 16 memory input positions, PixelMem uses 74.25% fewer input cost than full text and 93.98% less storage than rendered-text images under the same input budget. Its accuracy is 89.4%, only 3.1 percentage points below full text. In a separate comparison with memories stored as vectors, it reaches 76.2% accuracy with four memory input positions, exceeding the vector baseline's 48.5% by 27.7 percentage points. These results suggest that small pixel images can preserve information useful for answering questions while requiring few model input positions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.