acceptodds
Under review as a conference paper at ICLR 2027

AsynCompact: Efficient Context Compaction for Long-Horizon LLM agents

Abstract

Context compaction supports long-horizon LLM agents but can interrupt execution during summary generation and KV-cache reconstruction. We present a compaction framework that reduces these interruptions through asynchronous execution and optional KV grafting. The agent continues its task while a background branch generates a summary, then adopts the compacted history at a safe step boundary. For self-hosted inference, KV grafting further reuses cached states for the summary and retained context, with positional rebasing, instead of fully recomputing the compacted context. This reuse is approximate and requires no model retraining. Experiments on three benchmarks show broadly comparable task performance. Controlled boundary replays show that asynchronous compaction reduces mean explicit summary blocking by –. In a separate Qwen3.8-27B replay with one active request per replica, KV grafting reduces mean post-compaction resumption latency by , including graft-specific preparation costs. These results demonstrate reduced waiting at compaction boundaries, supporting more responsive agent execution without relying on consistent full-task speedups.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.