MACC: Multi-Aspect Commitment Control for Drift-Aware Memory Management in Long-Horizon LLM Agents
Abstract
Long-horizon dialogue agents built on large language models (LLMs) must preserve still-active user commitments while operating under limited context budgets. We study context drift, unintended divergence from such commitments, and ask whether efficient memory editing also preserves them. Multi-Aspect Commitment Control (MACC) scores historical segments with several projection-based commitment alignments and learns to retain, summarize, or prune them, using a response-consistency proxy computed by an encoder separate from the generator. On the LoCoMo long-term conversational-memory benchmark and on Blended Skill Talk, with a shared Qwen2.5-1.5B-Instruct generator, MACC reaches 62.9% judge-rated acceptability on LoCoMo at 0.092 of the Full-History token cost. A No-History control reaches 60.4% at 0.087, so acceptability alone does not establish memory use. On a human-labelled diagnostic of 60 distance, distractor, and revision probes, three heuristic memory methods fail on 55–67% of probes, and a post-hoc evaluation of one MACC checkpoint fails on 78% with no dialogue history retained. These results expose a gap between task–cost performance and commitment fidelity and motivate explicit commitment testing alongside task–cost evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.