acceptodds
Under review as a conference paper at ICLR 2027

What Still Matters: Compressing Long Context via Set-Conditioned Marginal Value

Abstract

Long-context inference is essential to many applications of large language models (LLMs), yet processing lengthy inputs remains computationally expensive and often suffers from substantial information redundancy. Existing context compression methods alleviate this burden by reducing input length, and recent approaches further assign non-uniform compression rates according to segment-wise importance. In this work, we argue that the value of preserving a segment at high fidelity is inherently set-dependent: a segment that appears highly relevant in isolation may contribute little once its evidence is already covered, while a less relevant but complementary segment can provide greater marginal utility. To capture this dependency, we introduce ReValue a set-conditioned context compression framework that continually reassesses what still matters as the high-fidelity set evolves. ReValue first derives compact query-aware representations for context segments, and then constructs a Determinantal Point Process (DPP)-based set objective that jointly captures query-conditioned quality and multi-scale semantic overlap. Rather than relying on fixed segment-wise scores, each candidate segment is evaluated by its marginal contribution relative to the current high-fidelity set, allowing redundant evidence to be progressively discounted while complementary information remains valuable. The resulting hybrid context preserves full token-level representations for high-value segments and compact representations for the rest, further supported by a representation diversity regularizer. Experiments demonstrate consistent gains across backbones, compression ratios, and benchmarks, including up to 11.04 points in average F1 and 5.01 faster first-token latency.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.