RegimeCouncil: Paired-Evidence Routing for Memory Conflict Governance in LLM Agents
Abstract
Language agents increasingly rely on long-term memory to sustain extended interactions with their users, yet stored facts become stale as the world changes, and these agents often learn that a fact was wrong only through delayed feedback, after already acting on that fact. A central question is how far a single correction should propagate: should it govern only later uses of the same fact, or also how much the agent trusts memory on other facts? Existing systems typically adopt one of these two policies at design time. We formalize this choice as a sequential decision problem—memory conflict governance—in which memory validity is hidden at decision time and outcomes are revealed with delay, and prove that neither policy is uniformly sufficient. We then propose RegimeCouncil, which retains both policies and selects between them using delayed paired outcomes, without retraining or assuming a fixed dominance order. Two advisors score the advantage of using memory over a no-memory anchor: Global pools evidence across stored facts, while Local uses only the queried fact. A Meta Governor defaults to Local and switches to Global only when Global sustains a positive recent margin; under delayed feedback, this switch is provably reached within a bounded number of decisions. On delayed-feedback streams, RegimeCouncil outperforms the better fixed-scale baseline by – points on disagreement episodes. The same frozen router transfers to naturally occurring drift on TempLAMA and TAQA. Ablations confirm that every component is important: removing each of Global, Local, and the Meta Governor lowers accuracy by – points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.