acceptodds
Under review as a conference paper at ICLR 2027

COMPACT-Bench: One Budget Axis for Memory Compaction in LLMs and Agents

Abstract

Methods that reduce how much of its context a language model keeps in memory are developed in four separate literatures, and each reports results in its own unit. KV-cache compression reports memory saved at fixed perplexity, prompt compression reports tokens saved at fixed accuracy, sub-quadratic architectures report recall at a fixed state size, and agent-memory systems report benchmark scores without stating how much they store. All of these methods spend the same resource, yet their numbers cannot be compared. We introduce COMPACT-Bench, which measures every method on one axis, the bytes of retained state per token of history (BPT), and asks four questions at matched budgets: how much accuracy a method keeps, whether exact or lossy retention does better, whether the system can tell what it discarded, and whether its stated confidence reflects what its memory still supports. Every self-report is checked against uncompacted controls in which the fact is present or absent. On 21 settings of 16 open models from 0.5B to 24B parameters we find four results. First, the literatures disagree once every stored byte is counted. On Qwen3-1.7B a 2-bit KV cache keeps 0.98 accuracy at 22% of the full cache, where eviction, prompt compression, and summarization keep at most 0.56 within 30%, while on Qwen2.5-1.5B the same quantizer fails below 8 bits. Second, tolerance to eviction depends on model generation and not on size. In five same-family pairs from 0.5B to 14B, and across all 21 settings, every model released from November 2024 onwards loses half its accuracy at 33–69% of its budget, against 76–86% for every earlier model. Third, content-based scorers fall below random eviction only because they cannot see the question, and letting them see it recovers up to 0.98 in accuracy at the same budget. Fourth, eviction that scatters its losses leaves fragments of a fact that the model takes for the fact itself. On the models whose yes/no self-report passes both controls, SnapKV's failures are reported as still answerable in up to 74% of cases, against at most 11% for StreamingLLM, which removes facts whole, and on the models whose stated confidence falls for an absent fact, SnapKV's failures carry higher confidence than a fact that was never there. We release the benchmark and a reference implementation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.