AdaMem: Adaptive Memory Organization for Scalable Long-Context Reasoning
Abstract
Large language models (LLMs) face fundamental scalability challenges when reasoning over extremely long contexts, where processing the full sequence incurs prohibitive computational and memory costs. Existing context compression methods alleviate this burden by transforming long contexts into compact memory representations, but they typically rely on fixed segmentation and uniform compression strategies, leading to semantic fragmentation and inefficient information preservation. To address these limitations, we propose AdaMem, an adaptive memory organization framework that enables scalable long-context reasoning by dynamically organizing, allocating, and preserving contextual information. AdaMem first identifies semantically coherent units through boundary-aware chunking, avoiding information fragmentation caused by rigid partitioning. It then estimates the information novelty of each chunk and adaptively allocates memory budgets, allowing critical evidence to retain richer representations while compressing redundant content aggressively. Finally, AdaMem accumulates chunk-level memories and introduces selective replay to recover original contexts only when compressed memories are insufficient. Extensive experiments on long-context question answering benchmarks up to 1M tokens demonstrate that AdaMem consistently outperforms state-of-the-art long-context models, achieving superior reasoning accuracy with substantially improved efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.