acceptodds
Under review as a conference paper at ICLR 2027

SearchTrimmer: Search-Selective KV Eviction and Loop Gating for Efficient Deep Research Agents

Abstract

Deep research agents incur costs from both an expanding history and repeated rounds of model inference and tool use. The growing history increases KV-cache memory, while compressing it can add rounds as agents attempt to recover missing observations. We propose SearchTrimmer, a training-free framework that addresses both costs through two components motivated by an analysis of agent trajectories. First, search observations occupy most of the history while containing little unique evidence and remaining inexpensive to recover. We therefore propose Search-Selective KV Eviction, which retains browse observations and reasoning while keeping only the most recent search observations in the KV cache. Second, additional search activity often repeats earlier queries. We distinguish two types of re-search: Restore, which leads to a newly retained browse observation, and Loop, which does not. Since Loops account for most of the increase, we propose Loop Gating, which limits their cumulative count and requests a final answer when the budget is reached, while leaving Restores unpenalized. Neither component requires retraining, an additional model, or access to attention scores. Across three benchmarks, SearchTrimmer reduces average KV-cache usage by – on WebExplorer-8B and – on WebSailor-32B, with maximum KV-cache reductions of – and –, respectively, while matching or exceeding the uncompressed agents' avg@3 accuracy and keeping round counts within approximately of their baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.