Training with Recoverable Context Improves Agents Beyond Recall at Inference
Abstract
Context management shapes not only what an agent can access at inference time but also what it can learn during reinforcement learning. When truncation causes all rollouts in a group to receive the same reward, group-relative advantage estimation removes that group’s learning signal. We ask whether making truncated observations recoverable during training can improve the learned policy beyond the immediate benefit of recall at inference. Our method stores oversized observations outside the active context, replaces them with compact references, and allows the policy to retrieve them under explicit budgets, while treating recalled text strictly as observation rather than a prediction target. Across four model scales, recall-enabled training improves mathematical accuracy even when recall is disabled at evaluation, indicating that recoverable context changes what the policy learns, not just what it can access. On long-context tasks, it also substantially increases answer coverage. Controlled experiments rule out unprompted recall behavior and generic long-context reasoning as sufficient explanations. These findings establish context management as a training intervention, not merely an inference-time design choice, and highlight the importance of separating its learning effects from its retrieval benefits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.