acceptodds
Under review as a conference paper at ICLR 2027

Learn to Remember: Geometric Memory for Inference-Time Self-Improvement in Language Models

Abstract

Large language models (LLMs) are used in a stateless way, discarding what they learn from one query to the next. Memory-augmented LLMs adapt at inference time, but existing systems either inject the entire memory into every prompt, so cost grows with it, or retrieve by similarity under a fixed encoder that never learns whether an entry helped. We propose Learn to Remember (LeRe), in which a frozen LLM cycles through plan, solve, and curate roles and learns which memory to retrieve rather than how much to store. A Planner rewrites each query into a structured retrieval key; the Contrastive Contextual Memory Encoder (CCME), lightweight heads on a frozen encoder, is trained online from the Curator's attribution of which entries helped, so retrieval reflects measured utility; and a Guard–Consolidate–Maintain (GCM) gate rejects answer leakage, merges near-duplicates, and quarantines unreliable entries. LeRe learns without ground-truth labels or weight updates. Across twelve mathematical, scientific, and multimodal streams and three backbones, LeRe improves average accuracy over a no-memory baseline by 6.9%, 9.0%, and 4.6%, exceeding the strongest memory-based baseline on every backbone and gaining up to 24.7% on competition mathematics, while its per-query token cost stays flat as memory grows, at near-lowest cost. These consistent gains suggest that selecting proven memory, rather than accumulating more, offers a scalable, efficient, and effective way to support inference-time adaptation in LLMs. Our code is available in an anonymous repository at https://anonymous.4open.science/r/learn-to-remember-AD37.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.