Governing Agent Self-Improvement Across Timescales
Abstract
LLM agents can self-improve during deployment through mechanisms operating at different timescales, such as within-task refinement and cross-task memory. These mechanisms can improve performance, but each additional refinement round or memory update also incurs inference cost, raising a fundamental question: when is continued self-improvement across timescales still worth the additional cost? We study this question through ReMo, a controlled dual-timescale framework for investigating within-task refinement and cross-task memory under a common agent workflow. Across general agentic and financial tasks, model sizes, and model families, both mechanisms improve performance and provide complementary gains when combined. However, their marginal value diminishes as self-improvement continues. Within tasks, additional refinement rounds can add computation with little further benefit; across tasks, memory can continue to grow after further updates provide little measurable improvement. These results show that fixed self-improvement policies can continue consuming inference computation after their marginal value has diminished. Motivated by this observation, we introduce AutoGovern, a lightweight online strategy that automatically governs whether refinement and memory updating should continue during execution. Across AppWorld and Formula, AutoGovern reduces token cost by 38% on average and by up to 72% while maintaining comparable task performance. Our results suggest treating agent self-improvement as inference computation whose marginal value should be governed over time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.