BudgetWeaver: Learning the Downstream Value of Agent Reads and Memory Mutations under Shared Budgets
Abstract
Every optional-memory action changes two things at once: the current context and the evidence–budget state inherited by later decisions. A write can preserve verified progress, an eviction can remove a stale hypothesis, and a summary can make later evidence affordable. Controllers that score each operation only for its immediate use miss this intertemporal choice when every read and mutation spends the same token–latency budget. BudgetWeaver learns the downstream value of a seven-operation sequence over NOOP, retrieval, reranking, graph traversal, summarization, writing, and eviction. Its policy combines remaining budget and memory state with paired read-versus-NOOP endpoint outcomes, support-aware offline value refinement, predicted-cost admission, and realized-cost accounting. We isolate the controller by fixing prompts, candidates, operators, retrievers, stores, schemas, and guards across all matched policies. At the predeclared medium operating point, BudgetWeaver improves native success over tuned full-interface non-RL control by 1.8–2.6 points on LOCOMO v2, SWE-bench Lite v1.1, and OdysseyBench v1.3; every paired interval excludes zero and all three tests remain significant after Holm correction. Relative to always-retrieve, success rises by 3.7–4.9 points while paired full-run token use falls by 4.1k–5.4k and end-to-end p95 latency falls by 1.6–2.7 seconds. Removing persistent mutations costs 1.3–2.8 points, with the largest descriptive effects in contradiction-heavy and longest-horizon episodes, and separately fitted Llama-3.1-8B-Instruct controllers retain a 1.7–2.6-point lead over tuned non-RL control. These results make downstream memory-state value an actionable control signal: coordinating reads and mutations improves success while reducing complete-run resource use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.