ANCHOR: AMORTIZED KV-CACHE COMPACTION FOR PROCEDURAL SKILLS
Abstract
LLM agents increasingly rely on long procedural skills: prompts packaging tool schemas and invocation instructions that consume extensive context on every request. KV-cache compaction promises to reclaim this budget by emitting a compacted KV prefix for unseen skills in a single forward pass, yet existing methods perform poorly on procedural skills, producing missing or inaccurate tool calls. We identify a key training mechanism behind this, the anchoring effect: when knowledge-preservation (reconstruction) and instruction-following (tool-calling) objectives are trained sequentially, the second stage erases most of what the first stage stored, leading to sub-optimal final tool-call accuracy. We therefore present Amortized Neural KV Compaction with Held-On Reconstruction, an emitter network trained with on-policy distillation anchored by a low-weight reconstruction objective that keeps stored knowledge intact while improving tool-call fidelity. We also find that evaluating skill compaction is itself an open problem: on some existing agentic benchmarks, strong base models score nearly as well without the skill in context, leaving no headroom to measure what a compacted prefix preserves. We therefore propose DomainBench: 716 tasks from 100 procedural skills, answerable only with the skill in context and graded on executable tool calls at three nested granularities of tool name, argument keys, and argument values. At 4x compaction ANCHOR recovers 65.1% and 58.9% of full-context tool-use fidelity on two LLM backbones against 5.5–45.0% for four prior methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.