Beyond Token Pruning: Prompt Compression via Minimum Description Length and Semantic Message Anchors
Abstract
Prompt compression accelerates long-context inference by condensing extensive context into concise core messages. Existing approaches follow an extractive score-and-drop paradigm, scoring tokens or text spans via local importance heuristics and discarding low-ranked elements. Hard deletion permanently discards distributed background context, failing on tasks requiring global counts or bounds, while token pruning destroys syntactic integrity and narrative topology. We reframe prompt compression as a coding problem rather than a filtering problem. Under the Minimum Description Length (MDL) principle, a compressed prompt is a two-part code: intact evidence units answer the query directly, while compact sufficient statistics summarize omitted background context. This view yields a select-and-aggregate paradigm beyond pruning, which we instantiate as emantic essage nchors (SMA), an efficient, training-free, and state-agnostic framework requiring no reader internal states or gradient updates. SMA segments text into grammatically self-contained propositional units, greedily packs query-salient anchors while preserving relative document topology, and compiles unselected background spans into compact sufficient statistics. SMA outperforms the strongest evaluated training-free compressor by up to 9.34 points on LongBench, accelerates prefill by up to 7.55, and reduces peak KV cache by 10.9, advancing the empirical Quality-Latency Pareto frontier.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.