Reversible Memory Management for LLM Agents with Support for Token-Level Supervision
Abstract
Summarizing agent histories can discard details needed later, while recalling historical text repeats prefill computation. We introduce RETURN, a framework for reversible context management that keeps active context compact while retaining historical KV states. It restores selected turns with updated positions, without re-prefilling their text. An ephemeral probe selects history from turn digests, then rolls back its tokens and execution state to keep memory management outside the task trace. This enables supervision from a full-context teacher without management labels. Our backend offloads inactive KV to CPU, while path-aware parallel replay reconstructs recorded history selections and position mappings and propagates gradients through reused states. Across memory and agentic benchmarks, RETURN generally delivers stronger and more consistent performance than text-recall baselines without additional training. It also improves accuracy beyond the native context window. Despite probing overhead with short outputs, reduced GPU KV memory enables higher concurrency and improves serving throughput. On-policy distillation with parallel replay improves task performance, demonstrating compatibility with token-level supervision on formal assistant tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.