TraSURE: Exploiting Cross-trajectory Recurrence to Accelerate Information-seeking RL Rollouts over Shared Corpora
Abstract
Training agents with reinforcement learning (RL) to answer questions over shared corpora, such as enterprise knowledge bases, incurs substantial rollout cost. We identify cross-trajectory recurrence in this setting: trajectories for the same or different queries retrieve the same external items under different contexts, causing repeated prefill that group-based sampling further amplifies. However, this recurrence remains unexplored for rollout acceleration. We present TraSURE, which exploits this recurrence to accelerate RL rollouts over shared corpora. (1) For prefill, we use position-independent caching (PIC) to share recurring-item KV states across trajectories and avoid redundant prefills. (2) Building on this caching layer, we design demand-aware rollout scheduling to coordinate cache admission and retention using current and future item demand. (3) Beyond prefill, we use a next-outcome posterior to select historical continuations as drafts for posterior-guided speculative decoding. Across six models and four information-seeking workloads over shared corpora, TraSURE delivers up to speedup over standard rollouts, with all task scores within 1 pp of full-causal counterparts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.