CIT-Mem: Counterfactual Intervention Training for Evidence-Grounded Memory Policies in Long-Horizon LLM Agents
Abstract
Long-horizon LLM agents rely on memory to interpret conversational history, yet most systems optimize semantic relevance rather than whether a record remains valid, directly supports the query, and provides sufficient evidence. A correct answer may still reflect parametric knowledge or a shortcut; inference-time interventions can reveal this dependence but do not teach a policy how to use evidence. We introduce CIT-Mem (Counterfactual Intervention Training for Memory), a framework that jointly trains invariance when the role of decisive evidence is preserved, sensitivity when that role changes, and abstention when no valid and sufficient evidence remains. It constructs validated non-causal twins that preserve decisive evidence and causal twins that remove, replace, contradict, or update it. Validation checks structural consistency, evidence provenance, recoverability, and answer direction, discarding unreliable transformations before training. Aligned policy objectives directly supervise memory operations, retrieval, and answerability, while adaptive replay prioritizes unresolved constraints with the backend and reader fixed. Experiments on LoCoMo-10 and LongMemEval-S show that CIT-Mem improves intervention consistency and abstention reliability while attaining the best overall answer quality across three frozen readers. On LoCoMo-10 with Qwen3-8B, it improves all-query F1 by 14.3% and BLEU-1 by 17.2% relative to the strongest baseline while using the smallest retrieved-context budget. It also achieves the best overall judged accuracy on LongMemEval-S, demonstrating more effective evidence-grounded memory selection in long conversations. Our code can be found at https://anonymous.4open.science/r/cit-mem-2B43.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.