acceptodds
Under review as a conference paper at ICLR 2027

ARKA: TYPED MEMORY ACCESS FOR EFFICIENT LONG-CONTEXT INFERENCE

Abstract

Long-context LLM agents accumulate dialogue histories, task traces, exemplars, and evolving facts, and full-context prompting exposes this growing state to the backbone LLM, making the KV cache expand with the interaction history. Full-context prompting repeatedly prefills irrelevant history, while prompt compression operates after long-context material has been gathered and typically requires an additional model pass. We propose ARKA, an LLM-free lightweight framework that treats agent memory as query-conditioned working-context assembly. ARKA stores histories and typed memory units outside the backbone LLM, uses a lightweight router to decide how memory should be accessed for each query, assembles compact evidence on CPU, and calls the backbone LLM at most once for generation. Our information- theoretic analysis establishes sufficient conditions for compact working contexts to preserve answer-relevant information and shows that discarding required source relations can prevent answer recovery even when relevant words remain. The LLM-free routing framework is evaluated on Llama-3.1-8B, Qwen3-8B, and Qwen3.5-9B backbones without backbone-specific retraining. Across MemoryAgentBench, LongBench-5, LoCoMo, and MemBench, ARKA outperforms the mBERT-based LLMLingua-2-small compressor in all 11 comparable benchmark–backbone settings, with an average gain of 9.9 points and a median 6.4× speedup in recorded processing time. In a history- growth stress test, full-context prompting runs out of memory at 50× scale on an A100 80 GB GPU, while ARKA reaches 100× with only a few hundred backbone input tokens. These results show that scalable agent memory should be built by learning which memory evidence to activate and which access policy to use for each query, rather than by extending or compressing a monolithic prompt.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.