Learning What to Refresh: Counterfactual Observation Selection for Tool-Using Agents
Abstract
Tool-using language agents often act on cached observations collected earlier in the trajectory, even though the external world may have changed in the meantime. This induces a belief-repair problem: the agent's internal state remains anchored to historical evidence whose relevance depends on the current task objective and the evolving environment. Re-querying every source is prohibitively expensive, while observation age alone is not a reliable indicator of whether a refresh will improve the next decision. We study task-aware observation refresh: selecting a subset of previously observed sources to revisit before the agent continues execution. Rather than treating staleness as a binary signal or a recency heuristic, we learn refresh decisions from counterfactual interventions at the same task checkpoint, comparing downstream task utility under different refresh sets while holding the agent state fixed. A set-conditioned critic estimates the joint value of candidate refresh sets, allowing the policy to account for complementary and redundant observations while balancing expected utility against acquisition cost. The controller operates on agent-visible history, cached observations, timestamps, and query costs, with the language-model backbone kept frozen. We evaluate the approach on a controlled synthetic benchmark with stale observations and task-dependent source interactions, and compare against no-refresh, always-refresh, recency-based refresh, and additive-value baselines. Our goal is to characterize when selective refresh improves the success–cost trade-off and whether the learned policy transfers to unseen task structures and environment dynamics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.