acceptodds
Under review as a conference paper at ICLR 2027

From Cache to Belief: Online Probabilistic Memory for Long-Context Language Models

Abstract

Language models that serve long documents, dialogues, and agents must reuse evidence that has left their attention window, and full attention over the whole history is costly. Compressed memories keep a bounded state instead, but store it as point estimates: a slot cannot tell whether it summarizes many consistent observations or a single noisy one, so how far to trust or overwrite it must be learned implicitly. We argue that bounded memory should maintain beliefs rather than points. BeliefMem is a lightweight adapter that keeps the history beyond the window of a frozen sliding-window model as a bank of Gaussian slot beliefs. One posterior responsibility decides where each evidence block is written and which slots are read; natural-parameter writes let accumulated precision resist overwriting, and a precision-weighted product of experts lets the read revise local evidence only as far as memory is more certain. Trained on 600 examples (27 minutes on one GPU at 7B) with 0.11–0.23% of the backbone parameters, BeliefMem improves six-task LongBench over the local model by 2.72, 4.05, and 2.65 points on Qwen2.5-3B/7B/14B, outperforms all nine Artificial Hippocampus Network (AHN) checkpoints at every scale, and transfers to LongBench v2 and InfiniteBench. A retrained point memory with the same interface trails it in every paired seed, and a causal probe shows that its benefit grows when the answer evidence lies beyond the window.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.