BCAR: Budgeted Concave Adaptive Retention for Query-Aware Context Compression in Long-Running Agents
Abstract
Long-running agents must answer under a fixed prompt budget even after relevant turns have been retrieved. Keeping the full pool is slow and noisy, and compressing it as one flat token stream often drops the evidence that the current query needs. We ormalize this decision as query-adaptive budgeted context allocation: each retrieved turn is ssigned a retention ratio under a global character budget. BCAR implements this with a train-free query–turn value score, chunked water filling quotas, and BEAVER hierarchical span selection (HSP) inside each quota. At a 10% keep ratio, BCAR reaches 64.8% on LongMemEval-S versus 57.4% for retrieve-all (+7.4 points), and outperforms the other compressors under the same retrieval and a shared character or ≈5800-token budget. On AgentLongBench the mean difference versus retrieve-all is +2.8 points at 32k and +2.5 at 64k; neither pooled comparison is significant, and the two 64k knowledge-intensive splits are slightly negative. Under the same quotas on LongMemEval-S, HSP with role-aware pan ordering is 7.0 points above prefix head truncation. Perceived latency (prep+TTFT) drops by 9.7–18.1×.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.