Task-Oriented Memory Compilation: Executable State Representations for Long-Context Language Models
Abstract
Answering questions over long interaction histories often requires tracking updated facts, following relations, and aggregating scattered evidence. Retrieval and compression methods supply selected text, which can leave these operations to the answering model, and typically rely on auxiliary neural models to produce it. We propose Task-Oriented Memory Compilation (TOMC), which instead performs these operations while building memory. TOMC computes state, relation, or count records from the history and packs them with selected source text under a fixed token budget. It is training-free, runs on CPU, and supplies plain text to black-box LLM APIs. On the BEAM long-term memory benchmark, with 100K–1M-token histories and three strong LLMs as API readers (GPT-5.1, Rednote preview, and DeepSeek 4.1 Flash), TOMC achieves higher overall mean scores than six retrieval, memory, and compression baselines. It also uses 32.4–36.6% fewer input tokens than the strongest baseline, LIGHT, and 11–17× less construction time than per-call neural compressors. At 500K and 1M tokens, TOMC scores higher in all 36 reader–baseline–length comparisons; its largest gains over BEAM-RAG and LIGHT are in temporal reasoning, which ablation links to the records. These results suggest that executing task operations during memory construction complements retrieval and offers an alternative to token compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.