When Agents Should Not Remember: Habitual Execution for Retrieval-Efficient Long-Running LLM Agents
Abstract
Long-running LLM agents often retrieve long-term memory even on stable, repeated workflows, which adds retrieval cost, stale context, and operational complexity. Reducing retrieval is not simple caching: the agent must decide, before retrieval, which workflows are SOP-stable, and must interrupt an automatic routine when policy, tool, or user-goal signals show that the context has changed. We propose habitual execution, a complementary control path. A Habitual Memory Layer compiles verified successful trajectories into guarded routines that can bypass explicit retrieval when the context is stable, and return control to retrieval when the routine no longer applies. On tool-using airline and retail customer-service tasks, HM-Gated RAG matches always-retrieve RAG in mean success over four trials (airline 0.667 vs. 0.675; retail 0.497 vs. 0.506; 95% CIs on the paired gap are [-8.3,+6.7] and [-6.9,+5.0] percentage points, both crossing zero) while cutting retrieval calls by about 59% and 51% (95% CIs [40.7%,76.1%] and [40.0%, 61.9%]). The saving requires offline task-type tags; keyword and no-tag routers do not produce it. The method is not a general way to reduce local tokens or latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.