When to Look Again: Learning from Refresh Feedback For LLM Agents
Abstract
Large language model (LLM) agents reuse information from tools, retrieval, and memory over many steps while the underlying sources keep changing, so an agent can act on an outdated copy of information it once retrieved correctly. With a small refresh budget per step, the agent must decide which copies can wait. We model independent source changes with weighted staleness costs and a shared refresh budget. The model is a restless bandit whose closed-form Whittle index sets the refresh priority and a Lagrangian lower bound evaluates the scheduler. The index reflects both the risk of staleness and how long a refresh keeps the copy fresh, i.e., faster change raises priority at short ages but lowers the priority ceiling. When rates are learned from refreshes, an upper confidence bound (UCB) on a source’s rate therefore lowers its priority and can delay the observations that would correct it. We instead maximize the index over plausible rates and rebuild the plausible set at every decision. For a two-source class, we prove that freezing the set can yield linear expected excess cost while rebuilding it keeps the excess bounded, and we give checkable conditions for bounded cost with multiple unknown sources. In numerical experiments, maximizing the index over plausible rates approaches the cost of a scheduler that knows the rates, while a UCB on the rate incurs persistent excess cost. With known rates, the index scheduler stays close to the lower bound across budgets and system sizes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.