SearchTrimmer: Search-Selective KV Eviction and Loop Gating for Efficient Deep Research Agents
Abstract
Deep research agents incur costs from both an expanding history and repeated rounds of model inference and tool use. The growing history increases KV-cache memory, while compressing it can add rounds as agents attempt to recover missing observations. We propose SearchTrimmer, a training-free framework that addresses both costs through two components motivated by an analysis of agent trajectories. First, search observations occupy most of the history while containing little unique evidence and remaining inexpensive to recover. We therefore propose Search-Selective KV Eviction, which retains browse observations and reasoning while keeping only the most recent search observations in the KV cache. Second, additional search activity often repeats earlier queries. We distinguish two types of re-search: Restore, which leads to a newly retained browse observation, and Loop, which does not. Since Loops account for most of the increase, we propose Loop Gating, which limits their cumulative count and requests a final answer when the budget is reached, while leaving Restores unpenalized. Neither component requires retraining, an additional model, or access to attention scores. Across three benchmarks, SearchTrimmer reduces average KV-cache usage by – on WebExplorer-8B and – on WebSailor-32B, with maximum KV-cache reductions of – and –, respectively, while matching or exceeding the uncompressed agents' avg@3 accuracy and keeping round counts within approximately of their baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.